How-To Guides¶
As large datasets of enumerated fingerprints and other information for cheminformatics might not fit into working memory, ᵡ-SOM supplies a dedicated HDF5 file layout and HDF5Dataset class. This HDF5Dataset class is compatible with the PyTorch DataLoader for random, millisecond latency, access into this on-disk storage.
Creating a HDF5Dataset file¶
ᵡ-SOM supplies a tool to generate the HDF5 files containing the fingerprints directly from text files of molecular data in different line notations, e.g. SMILES, INCHI, etc., or text files already containing enumerated fingerprint data. Other properties of the molecules can also be recorded for examination with the GUI.
We import HDF5Creator and either rdStyleFactory or CSVStyleFactory. Both need an rdMolGenerator to build the internal representation from line notation. rdFingerprintGenerator is specific to the direct generation of enumerated fingerprints.
from numpy import float32
from rdkit.Chem import MolFromSmiles, rdFingerprintGenerator
from chisom.io.datastore_creation import HDF5Creator
from chisom.io.datastore_factories import rdStyleFactory
generator = rdFingerprintGenerator.GetMorganGenerator
fingerprint_kwargs = {"fpSize": 1024, "radius": 2}
file_dict = {
"active": [
"tests/VDR/actives.smi",
],
"inactive": [
"tests/VDR/inactives.smi",
],
}
molgen = rdStyleFactory(
MolFromSmiles,
generator,
generator_kwargs=fingerprint_kwargs,
count_fingerprint=True,
)
file_creator = HDF5Creator(fingerprint_generator_factory=molgen)
leaf_map = {
"primary": (0, str),
"ID": (1, str, "na"),
"Activity": (2, int, "categorical"),
"MolWt": (3, float32, "continuous"),
"MolLogP": (4, float32, "continuous"),
"TPSA": (5, float32, "continuous"),
}
skip_lines defaults to 0 and sep to a tab.
file_creator.create(
file_dict,
"tests/VDR.h5",
leaf_map,
skip_lines=1,
sep="\t",
)
Training an ESOM on data in a HDF5Dataset using CUDA¶
To use the generated HDF5Dataset PyTorch DataLoader needs to be set up accoringly.
PyTorch is not installed with ᵡ-SOM
torch is not a runtime dependency of ᵡ-SOM — only this DataLoader workflow needs it. Install it separately with pip install torch. Training from a plain NumPy array works without it.
The whole guide needs these imports:
import numpy as np
from torch.utils.data import DataLoader
from chisom import Som, start_chisom_viewer
from chisom.io import HDF5Dataset
from chisom.utils import lattice_size
"active" subset defined previously available for training. If omitted, the whole dataset is used for training.
# Create a ChemDataset object that from a HDF5 file, that is compatible with the pytorch DataLoader
ds = HDF5Dataset("tests/VDR.h5", ["active"])
# Create a DataLoader object that will be used to train the SOM
dl = DataLoader(
ds,
batch_size=BATCHSIZE,
shuffle=True,
num_workers=4,
)
rows, columns = lattice_size(len(ds))
# Create a SOM object
# The high and low parameters should be chosen according to the dataset values, to decrease the training time
som = Som(
rows,
columns,
ds.fingerprint_length,
"cosine",
use_cuda=True,
low=ds.fingerprint_min,
high=ds.fingerprint_max,
)
# Train the SOM for all epochs in a single call
som.train(dl, EPOCHS, ALPHA, BATCHSIZE)
umatrix property once training has finished. Saving it to disk lets you reopen the map later with chisom view.
# Calculate the U-map for the current state of the notebook
umx = som.umatrix
np.save("tests/umx.npy", umx)
# Create a DataLoader object for prediction (no shuffling)
dl = DataLoader(
ds,
batch_size=1000,
shuffle=False,
num_workers=10,
)
# Predict the best matching units and quantization errors for all data points
bmus, qe = som.predict(dl)
np.save("tests/bmus.npy", bmus)
Measuring distances on a trained SOM¶
The U-Matrix shows where the map is stretched, but it does not by itself tell you how far apart two regions are along the map. ᵡ-SOM exposes the underlying U-Distance graph for that: an undirected, edge-weighted graph whose nodes are grid positions and whose edges connect toroidal grid neighbours, weighted by the high-dimensional distance between the corresponding codebook vectors.
Build the graph once from the trained SOM and reuse it — it is cached on the Som, but binding it to a local name keeps repeated queries obvious and cheap.
from chisom import u_distance
graph = som.u_graph
u_distance then computes shortest-path lengths through that graph. The target is always a single (row, column) position; the source argument accepts three forms.
# 1. Distance from every unit on the map to a reference position
all_distances = u_distance(graph, None, (12, 30))
# 2. Distance from one position
single = u_distance(graph, (4, 4), (12, 30))
# 3. Distance from a collection of positions, e.g. the BMUs of a set of hits
hits = bmus[labels == "active"]
hit_distances = u_distance(graph, hits, (12, 30))
Each call returns a dict mapping the queried source node — as an (int, int) tuple — to its u-distance. Passing None yields one entry per node in the graph, which is the form to use when you want a full distance field to plot over the map.
Results are keyed by position, so duplicates collapse
Because the result is a dictionary keyed by grid position, passing a collection of BMUs returns one entry per distinct position. Several datapoints sharing a BMU yield a single entry, so the result can be shorter than the collection you passed in.
Build the graph once
Every call to u_distance runs a shortest-path search over the graph you hand it. Constructing the graph is the expensive part, so build it once and pass the same object to all your queries rather than accessing som.u_graph inside a loop over BMU pairs.
Rendering a figure without a display¶
The interactive viewer needs a graphical display. On a headless machine — a compute node, a CI job, an SSH session without X forwarding — use plot_som instead, which renders the same style of map through matplotlib and can write straight to a file. It is also the way to generate figures in bulk or from a script; if you already have the viewer open, it can export the map directly.
from chisom import plot_som
fig = plot_som(
som.umatrix,
bmu_coordinates=bmus,
data=ds,
color_by="Activity",
save_as="som_plot.png",
)
data accepts either a HDF5Dataset or a plain pandas.DataFrame, and color_by names one of its columns. Whether that column is treated as categorical or continuous is inferred, but you can force it with categorical=True/False. For a SOM trained with save_progress, layer selects which epoch's U-Matrix to draw (default -1, the last).
The figure sizes itself from the lattice, so non-square and very large maps come out with sensible proportions instead of being squeezed into a default canvas. The relevant knobs:
| Parameter | Default | Effect |
|---|---|---|
cell_size_in |
0.2 |
Inches per grid cell — the main size control |
chrome_scale |
1.0 |
Scales legend, colorbar and label sizes together |
marker_cell_fraction |
0.7 |
BMU marker diameter as a fraction of one grid cell |
legend_ncol |
auto | Columns in the categorical legend; derived from the category count if omitted |
legend_label_maxlen |
24 |
Long category labels are elided to this length |
Passing figsize or marker_size explicitly opts back out of the grid-derived sizing. Passing ax draws into an existing axes instead of creating a figure.
Using less cores than available¶
When ᵡ-SOM should use less core than are currently available on the machine, the desired number of cores to use (n_cores) can be set via numba, either programmatically
from numba import set_num_threads
set_num_threads(n_cores)
or via an environment variable
export NUMBA_NUM_THREADS=n_cores