Library Reference¶
This page documents the public Python API. Everything under chisom._core and
chisom._interface, and any module whose name begins with an underscore, is internal and
may change at any release — see Limitations.
The chisom command line interface is documented in The Viewer
rather than here, since its surface is its arguments rather than its functions.
chisom ¶
Classes:
-
Som–Main Class to create and train a Self-Organizing Map
Som ¶
Som(
rows: int,
columns: int,
features: int,
vector_distance: str = "euclidean",
map_distance: str = "euclidean_toroid",
neighborhood_kernel: str = "gaussian",
use_cuda: bool = False,
use_local_neighborhood: bool = False,
use_fastmath: bool = True,
save_progress: Optional[str] = None,
low: float = 0.0,
high: float = 1.0,
seed: Optional[int] = None,
)
Main Class to create and train a Self-Organizing Map
Parameters:
-
rows(int) –Number of rows of neurons.
-
columns(int) –Number of columns of neurons.
-
features(int) –Numbers of features to the data / weights of each neuron.
-
vector_distance(str, default:'euclidean') –Distance used in original data space, by default "euclidean". Possible values: "euclidean", "manhattan", "cosine"
-
map_distance(str, default:'euclidean_toroid') –Distance used in map space, by default "euclidean_toroid". Possbile values: "euclidean", "manhattan", "euclidean_toroid", "manhattan_toroid"
-
neighborhood_kernel(str, default:'gaussian') –Shape of the neighborhood kernel, by default "gaussian".
-
use_cuda(bool, default:False) –If True, CUDA accelleration is used. Needs numba-cuda-mlir, installed via the
cu12orcu13extra. By default False. -
use_local_neighborhood(bool, default:False) –Sets a hard neighborhood cutoff, by default False. Only used on CPU. Significantly increases performance at cost of numerical accuracy.
-
use_fastmath(bool, default:True) –Slightly decrease numerical accuracy to increase performance, by default True.
-
save_progress(Optional[str], default:None) –Saves codebook and U-Matrix to the given location if set, by default None. Usefull if long running computations crash / time out.
-
low(float, default:0.0) –Lower bound for codebook initialization, by default 0.0.
-
high(float, default:1.0) –Upper bound for codebook initialization, by default 1.0.
-
seed(Optional[int], default:None) –Randomness seed for replicability, by default None.
Raises:
-
ValueError–If the map dimensions are less than 1x1.
-
ValueError–If the number of features is less than 2.
-
ImportError–If CUDA is requested but not available.
-
ValueError–If the vector distance norm is not one of the supported norms.
-
ValueError–If the map distance norm is not one of the supported norms.
-
ValueError–If the neighborhood kernel is not one of the supported kernels.
Methods:
-
predict–Return the positions of the BMU for a dataset
-
train–Train the SOM with the given data for a number of epochs.
-
train_batch–Manually train the SOM with a single batch of data.
Attributes:
-
u_graph(Graph) –The U-Distance graph for the current codebook.
-
umatrix(UMatrix) –The U-Matrix for the SOM, as a 3D array of shape
u_graph
property
¶
u_graph: Graph
The U-Distance graph for the current codebook.
Undirected, edge-weighted networkx graph whose nodes
are grid positions (row, column) and whose edges connect toroidal
grid neighbors, weighted by the high-dimensional distance between
the corresponding codebook vectors, the same per-neighbor
distances used to build the U-Matrix. The graph is cached, but bind
it to a local name and reuse that for repeated
chisom.analysis.u_distance (or future ranking) queries.
Returns:
-
Graph–The U-Distance graph for the current codebook.
umatrix
property
¶
umatrix: UMatrix
The U-Matrix for the SOM, as a 3D array of shape (n_layers, rows, columns).
Computed and cached after training. When save_progress is set,
one layer is stored per epoch (training history); otherwise a single
final layer is stored. Accessing this before training computes a
single layer from the current (untrained) codebook on demand.
Returns:
-
NDArray[float16]–The U-Matrix, shape (n_layers, rows, columns).
predict ¶
predict(
data: NDArray | DataLoader,
) -> Tuple[NDArray[np.uint16], NDArray[np.float32]]
Return the positions of the BMU for a dataset
Parameters:
-
data(NDArray | DataLoader) –Dataset to find the BMUs for.
Returns:
-
NDArray[uint16]–The BMUs for the data.
-
NDArray[float32]–The Quantization Error
Raises:
-
TypeError–Error if the data format is not known
train ¶
train(
data: NDArray | DataLoader,
epochs: int,
alpha: float,
batchsize: int = 1,
shuffle: bool = True,
sigma: Optional[int] = None,
alpha_decay: str = "linear",
sigma_decay: str = "exponential",
alpha_end: float = 0.01,
sigma_end: int = 1,
) -> None
Train the SOM with the given data for a number of epochs.
Runs the full training loop internally, including per-step decay of
alpha and sigma and (if save_progress was set on
initialization) periodic U-Matrix/codebook checkpointing once per
epoch. Progress is reported via tqdm progress bars for both the
epoch loop and the per-epoch batch loop.
Parameters:
-
data(NDArray | DataLoader) –The data to train the SOM with. If a DataLoader is used, it is iterated batch by batch as configured on the DataLoader itself. If a numpy array is used, it is split into batches of size
batchsize. -
epochs(int) –Number of epochs to train for.
-
alpha(float) –The initial learning rate. Decays to
alpha_endover the course of training according toalpha_decay. -
batchsize(int, default:1) –Number of data points per batch. Only used when
datais a numpy array (ignored for DataLoader input, which defines its own batching), by default 1. -
shuffle(bool, default:True) –If True, the order of batches is reshuffled at the start of each epoch. Only applies when
datais a numpy array; DataLoader shuffling is controlled by the DataLoader itself. By default True. -
sigma(Optional[int], default:None) –The initial neighborhood radius. Must be greater than 0 if given. Defaults to None, in which case it is set to half the smaller of the map's row/column dimensions.
-
alpha_decay(str, default:'linear') –The decay schedule for
alpha, one of "linear" or "exponential", by default "linear". -
sigma_decay(str, default:'exponential') –The decay schedule for
sigma, one of "linear" or "exponential", by default "exponential". -
alpha_end(float, default:0.01) –The value
alphadecays towards by the end of training, by default 0.01. -
sigma_end(int, default:1) –The value
sigmadecays towards by the end of training, by default 1.
Raises:
-
ValueError–If
sigmais given and is less than or equal tosigma_end. -
ValueError–If
alphais given and is less than or equal toalpha_end. -
ValueError–If
alpha_decayorsigma_decayis not "linear" or "exponential".
train_batch ¶
train_batch(
data: NDArray, sigma: int, alpha: float
) -> None
Manually train the SOM with a single batch of data.
Provides fine-grained control over training by allowing the caller
to set alpha and sigma explicitly for a single batch, bypassing
the automatic per-epoch decay scheduling used by train. Useful for
custom training loops or schedules not covered by train.
Parameters:
-
data(NDArray) –A single batch of vectors to train the SOM on.
-
sigma(int) –The neighborhood radius to use for this batch.
-
alpha(float) –The learning rate to use for this batch.
start_chisom_viewer ¶
start_chisom_viewer(
umatrix: Optional[NDArray] = None,
bmu_coordinates: Optional[NDArray] = None,
data: Optional[DatasetBase | DataFrame] = None,
structure_info_column: Optional[str] = None,
scaling_factor: int = 3,
)
Start the GUI interface
All arguments are optional; anything not passed here can be loaded from the viewer's File menu instead.
Parameters:
-
umatrix(Optional[NDArray], default:None) –U-Matrix of the SOM.
-
bmu_coordinates(Optional[NDArray], default:None) –Coordinates of the BMU to the data points.
-
data(Optional[DatasetBase | DataFrame], default:None) –Additional data to the data points. Will be renderd in the tabel view and used for coloring of BMUs.
-
structure_info_column(Optional[str], default:None) –With chemical dataset the column with this index supplies the SMILES to render the molecule, by default None.
-
scaling_factor(int, default:3) –Will scale the U-Matrix by this factor ands interpolation for an anti-aliased view, by default 3.
chisom.utils ¶
Utility tools for hyperparameter configuration for (emergent) SOMs
Functions:
-
decay_exponential–Calculate the exponential decay of a value towards a desire final value for use in the training process.
-
decay_linear–Calculate the linear decay of a value for use in the training process.
-
lattice_size–Returns a rectangular lettice for the given number of data points,
decay_exponential ¶
decay_exponential(
iteration: int,
initial_value: int | float,
end_value: Optional[int | float] = None,
total_iterations: Optional[int] = None,
decay: Optional[float] = None,
*args,
**kwargs,
) -> float
Calculate the exponential decay of a value towards a desire final value for use in the training process. Ether to use a fixed number of iterations or a decay factor. 'decay' takes precedence over 'total_iterations'.
Parameters:
-
iteration(int) –Current iteration.
-
initial_value(int | float) –Starting value of the decay.
-
end_value(Optional[int | float], default:None) –Desired final value, when not using decay by default None.
-
total_iterations(Optional[int], default:None) –Total number of desired iterations, by default None.
-
decay(Optional[float], default:None) –Decay rate, by default None.
Returns:
-
float–Value at iteration
iteration.
Raises:
-
ValueError–If neither
total_iterationsnordecayis provided, or if bothend_valueanddecayare provided.
decay_linear ¶
decay_linear(
iteration: int,
initial_value: int | float,
total_iterations: Optional[int] = None,
decay: Optional[float] = None,
*args,
**kwargs,
) -> float
Calculate the linear decay of a value for use in the training process. Ether to use a fixed number of iterations or a decay factor. 'decay' takes precedence over 'total_iterations'.
Parameters:
-
iteration(int) –Current iteration.
-
initial_value(int | float) –Starting value of the decay.
-
total_iterations(Optional[int], default:None) –Total number of desired iterations, by default None.
-
decay(Optional[float], default:None) –Decay rate, like m in y = -mx+b, by default None.
Returns:
-
float–Value at iteration
iteration.
Raises:
-
ValueError–If neither
total_iterationsnordecayis provided.
lattice_size ¶
lattice_size(
dataset_size: int, factor=3
) -> Tuple[int, int]
Returns a rectangular lettice for the given number of data points, as recommendet by Ultsch et al. for ESOMs.
Parameters:
-
dataset_size(int) –Number of data points in the dataset.
-
factor–Ration of neurons to data points, by default 3.
Returns:
-
Tuple[int, int]–Number of (rows, columns)
chisom.analysis ¶
Post-training analysis of a trained SOM: distance and (future) ranking
queries over the SOM's U-Distance graph (see chisom.Som.u_graph).
Functions:
-
u_distance–Compute the u-distance(s) from source position(s) to a target position.
u_distance ¶
u_distance(
graph: Graph,
source: Optional[
PositionLike | Collection[PositionLike]
],
target: PositionLike,
) -> dict[tuple[int, int], float]
Compute the u-distance(s) from source position(s) to a target position.
The u-distance is the shortest-path length between a source and
target through a U-Distance graph (see chisom.Som.u_graph),
where edges are weighted by the high-dimensional distance between
neighboring codebook vectors. Build the graph once via
chisom.Som.u_graph and reuse it across many u_distance calls
(e.g. for many BMU pairs) rather than rebuilding it per call.
Parameters:
-
graph(Graph) –U-Distance graph, as returned by
chisom.Som.u_graph. -
source(Optional[PositionLike | Collection[PositionLike]]) –Grid position(s) to compute the distance from: -
None: every node ingraph. - A single position (tuple/array of 2 ints): that one node. - A collection of positions: each given source node. -
target(PositionLike) –Grid position (row, column) to compute the distance to. Same accepted forms as a single
sourceposition.
Returns:
-
dict[tuple[int, int], float]–Mapping from each queried source node (as an int
(row, column)tuple) to its u-distance totarget. Has one entry per node ingraphwhensource is None, a single entry for a single source position, or one entry per given source for a collection thereof.
Datasets are loaded by suffix. chisom.io.load_dataset (and the viewer's File → Load data
dialog) accept .h5/.hdf5 for HDF5 stores, .csv for comma-separated text, .tsv/.txt
for tab-separated text, and .parquet/.pq for Parquet. The full tuple is available as
chisom.io.loading.DATASET_SUFFIXES.
chisom.io ¶
Classes and Functions for large on-disk data stores specific to cheminformatics
HDF5Creator ¶
HDF5Creator(
fingerprint_generator_factory: DataloaderFingerprintGeneratorFactory,
file_extensions: list[str] = [".txt", ".smi", ".csv"],
num_processes: int = mp.cpu_count() - 2,
chunk_size: int = 1000,
queue_size: int = 500,
)
Bases: StoreCreator
HDF5 Store Create for creating HDF5 file of cheminformatic datasets from file hierarchies.
Parameters:
-
fingerprint_generator_factory(DataloaderFingerprintGeneratorFactory) –Factory to use, depends on the input file type and content.
-
file_extensions(list[str], default:['.txt', '.smi', '.csv']) –File extentions to consider.
-
num_processes(int, default:cpu_count() - 2) –Number of processes used.
-
chunk_size(int, default:1000) –Size of chunks send to processes. Should not need configuration.
-
queue_size(int, default:500) –Size of queue of each process,. Should not need configuration.
create ¶
create(
file_hierarchy: FileList,
out_path: str,
leaf_map: LeafMap,
skip_lines: int = 0,
sep: str = "\t",
) -> None
Run creation of HDF5 storage file
Parameters:
-
file_hierarchy(FileList) –Dictionary of files to parse, see How-To Guides.
-
out_path(str) –Output path for the HDF5 files.
-
leaf_map(LeafMap) –Dictionary of file structure and datatypes, see How-To Guides.
-
skip_lines(int, default:0) –Number of lines to skip at the beginning of file, e.g. for headers, by default 0.
-
sep(str, default:'\t') –Column seperator, by default "\t".
CSVStyleFactory ¶
CSVStyleFactory(
mol_generator: MolGenerator,
fpStart: int,
fpSize: int,
dtype: type = np.float32,
)
Bases: DataloaderFingerprintGeneratorFactory
A generator that creates fingerprints from a row of data, not from a Mol object. With a similar interface to the RDKit fingerprint generator, but using a slice of the row instead of a Mol object.
Parameters:
-
mol_generator(MolGenerator) –The RDKit MolGenerator to use to read Molecules in the input file e.g. MolFromSmiles
-
fpStart(int) –Numerical column index where the fingerprints starts.
-
fpSize(int) –Length in number of columns of the CSV file of the fingerprint.
-
dtype(type, default:float32) –Datatype of the fingperint elements
rdStyleFactory ¶
rdStyleFactory(
mol_generator: MolGenerator,
fingerprint_generator: rdFingerprintGenerator,
generator_kwargs={},
count_fingerprint=False,
)
Bases: DataloaderFingerprintGeneratorFactory
This class is designed to provide a unified interface for generating molecular fingerprints for use in DataLoader creation. It can combine different types of input (SMILES, InChI, ...) and fingerprint generator (e.g., Morgan, RDKit, Feature Morgan ... ). Also returns a new instance of the real generator on each call to get_generator, so that the generator can be used in a multiprocessing context without issues.
Parameters:
-
mol_generator(MolGenerator) –The RDkit MolGenerator to use to read Molecules in the input file. e.g. MolFromSmiles
-
fingerprint_generator(rdFingerprintGenerator) –The rdFingerprintGenerator to use.
-
generator_kwargs(dict, default:{}) –Keyword argument to pass to the rdFingerprintGenerator.
-
count_fingerprint(bool, default:False) –If count fingerprints should be created where applicable.
HDF5Dataset ¶
HDF5Dataset(
filepath: str, group_subset: Optional[List[str]] = None
)
Bases: DatasetBase
Loads HDF5 Datasets for large cheminformatic datasets created with the supplied script. Adheres to the PyTorch Dataset interface for the use with the PyTorch DataLoader for milisecond on-disc access data access.
Parameters:
-
filepath(str) –Path to HDF5 file
-
group_subset(Optional[List[str]], default:None) –Subsets to use, by default None
dataset_column_names ¶
dataset_column_names(
data: Union[DatasetBase, DataFrame],
) -> List[str]
Return the column names of either dataset flavour.
Parameters:
-
data(Union[DatasetBase, DataFrame]) –A dataset as returned by
load_dataset.
Returns:
-
List[str]–The column names.
inspect_hdf5_groups ¶
inspect_hdf5_groups(
filepath: Union[str, Path],
) -> List[str]
Read the group names of an HDF5 store without loading the dataset.
Parameters:
-
filepath(Union[str, Path]) –Path to the HDF5 store.
Returns:
-
List[str]–The names of the groups at the root of the store, sorted like
HDF5Datasetsorts them.
load_bmu_coordinates ¶
load_bmu_coordinates(filepath: Union[str, Path]) -> NDArray
Load BMU coordinates from a .npy file.
Parameters:
-
filepath(Union[str, Path]) –Path to a NumPy array of shape
(n_datapoints, 2)holding the row/column coordinate of the best matching unit of every datapoint.
Returns:
-
NDArray–The BMU coordinates.
Raises:
-
ValueError–Raised if the array does not have the expected shape or dtype.
load_dataset ¶
load_dataset(
filepath: Union[str, Path],
group_subset: Optional[List[str]] = None,
) -> Union[DatasetBase, DataFrame]
Load a dataset of datapoint properties, dispatching on the file extension.
Parameters:
-
filepath(Union[str, Path]) –Path to the dataset. Supported are HDF5 stores created with
HDF5Creator(.h5,.hdf5), delimited text (.csv,.tsv,.txt) and Parquet (.parquet,.pq). -
group_subset(Optional[List[str]], default:None) –HDF5 only: the groups to include, by default all of them.
Returns:
-
Union[DatasetBase, DataFrame]–An
HDF5Datasetfor HDF5 stores, aDataFrameotherwise.
Raises:
-
ValueError–Raised if the file extension is not supported.
load_umatrix ¶
load_umatrix(filepath: Union[str, Path]) -> NDArray
Load a U-matrix from a .npy file.
Parameters:
-
filepath(Union[str, Path]) –Path to a NumPy array holding a 2D (single layer) or 3D (layered) U-matrix.
Returns:
-
NDArray–The U-matrix, always as a 3D array with the layer as first axis.
Raises:
-
ValueError–Raised if the file does not contain a 2D or 3D array.
plot_som ¶
plot_som(
umatrix: UMatrix,
bmu_coordinates: Optional[NDArray[uint16]] = None,
data: Optional[Union[DatasetBase, DataFrame]] = None,
color_by: Optional[str] = None,
*,
categorical: Optional[bool] = None,
category_colors: Optional[Mapping[Any, Any]] = None,
cmap: Union[str, Colormap] = "viridis",
umatrix_cmap: Optional[Union[str, Colormap]] = None,
alpha_scheme: Optional[str] = "excess_absolute",
layer: int = -1,
scaling_factor: int = 3,
marker_size: Optional[float] = None,
marker_cell_fraction: float = 0.7,
legend_ncol: Optional[int] = None,
legend_label_maxlen: int = 24,
chrome_scale: float = 1.0,
figsize: Optional[tuple[float, float]] = None,
cell_size_in: float = DEFAULT_CELL_SIZE_IN,
dpi: int = 150,
ax: Optional[Axes] = None,
save_as: Optional[Union[str, Path]] = None,
) -> Figure
Plot a trained SOM as a static matplotlib figure.
Draws the interpolated U-matrix as background image and, if BMU coordinates are given, the BMUs as scatter markers on top — optionally colored by a property of the underlying data. Reproduces the default look of the interactive viewer, but works headless (e.g. with the Agg backend) and can be written directly to file.
Sizing of the map image, the BMU markers, and the colorbar(s)/legend is
coupled by default: the figure is sized from the SOM's actual grid shape
(cell_size_in), the axes box is aspect-locked to that same grid shape
(rows / columns), and marker size is derived from how large a single
grid cell actually renders once that layout is resolved. This keeps the
map, its markers, and its chrome proportionate to each other regardless
of how big or how non-square a given SOM's lattice is. Pass explicit
figsize/marker_size to opt back into fixed, grid-independent sizing.
Parameters:
-
umatrix(UMatrix) –U-matrix as returned by the
Som.umatrixproperty. Either 2D (rows, columns) or 3D (layers, rows, columns). -
bmu_coordinates(NDArray[uint16], default:None) –(N, 2) array of (row, column) BMU coordinates as returned by
Som.predict(). If None, only the U-matrix is drawn. -
data(DatasetBase or DataFrame, default:None) –Data source holding the property values referenced by
color_by. Must have one entry per row inbmu_coordinates. -
color_by(str, default:None) –Name of the column in
datato color the BMUs by. If None, BMUs are drawn in plain black. -
categorical(bool, default:None) –Force categorical (True) or continuous (False) coloring. By default this is taken from the column properties of
data(for a plain DataFrame the heuristic of the viewer is used: at most 10 unique values within the first 100 rows counts as categorical). -
category_colors(Mapping, default:None) –Mapping of category value to matplotlib color, used for categorical coloring. By default colors are assigned from the
tab10/tab20palettes in sorted category order. -
cmap(str or Colormap, default:'viridis') –Matplotlib colormap for continuous property coloring.
-
umatrix_cmap(str or Colormap, default:None) –Matplotlib colormap for the U-matrix. Defaults to the Earth colormap of the viewer.
-
alpha_scheme((gini, excess_absolute, excess_relative), default:"gini") –Weighting scheme used to encode the dominance of the primary category as marker opacity (see
RatioWeighting). If None, markers are fully opaque. Ignored for continuous coloring. -
layer(int, default:-1) –Layer to display for a 3D U-matrix.
-
scaling_factor(int, default:3) –Upscaling factor for the U-matrix interpolation.
-
marker_size(float, default:None) –BMU marker diameter in points. By default (None), it is derived from the rendered size of one SOM grid cell (see
marker_cell_fraction) so markers stay proportionate to the map regardless of grid shape or figure size. Pass a float to pin an exact, grid-independent size instead. -
marker_cell_fraction(float, default:0.7) –When
marker_sizeis auto-derived, the fraction of one rendered grid cell's size (the smaller of its width/height in points) that the marker diameter should occupy. Ignored ifmarker_sizeis set. -
legend_ncol(int, default:None) –Number of columns for the categorical-coloring legend. By default (None), chosen automatically so the legend wraps into additional columns rather than growing arbitrarily tall for large category counts (see
_auto_legend_ncol). -
legend_label_maxlen(int, default:24) –Maximum rendered length of a categorical legend label before it is elided with "…", so a few very long category names can't blow up the legend's (and therefore the map's) reserved width.
-
chrome_scale(float, default:1.0) –Single multiplier applied together to the auto-derived marker size, the colorbar's
fraction/pad, and the legend/colorbar font size, and to the chrome width reserved in an auto-derivedfigsize. Use this to scale all non-map chrome up or down in one step (e.g. for a much larger or smallerfigsize) instead of tuning each piece of chrome separately. Does not affect an explicitly-passedmarker_sizeorfigsize. -
figsize(tuple of float, default:None) –Figure size in inches, passed to matplotlib. By default (None), derived from the SOM's grid shape via
cell_size_ininstead of matplotlib's own grid-independent default. -
cell_size_in(float, default:DEFAULT_CELL_SIZE_IN) –Inches per SOM grid cell, used to derive
figsizewhenfigsizeis not given. Ignored iffigsizeis set explicitly. -
dpi(int, default:150) –Figure resolution, passed to matplotlib.
-
ax(Axes, default:None) –Existing axes to draw into. If given,
figsize,dpi, andcell_size_inare ignored and the legend is placed inside the axes. -
save_as(str or Path, default:None) –If given, the figure is saved to this path; the format is inferred from the suffix (e.g.
.pdf).
Returns:
-
Figure–The matplotlib figure, for further customization.