Skip to content

Library Reference

This page documents the public Python API. Everything under chisom._core and chisom._interface, and any module whose name begins with an underscore, is internal and may change at any release — see Limitations.

The chisom command line interface is documented in The Viewer rather than here, since its surface is its arguments rather than its functions.

chisom

Classes:

  • Som –

    Main Class to create and train a Self-Organizing Map

Som

Som(
    rows: int,
    columns: int,
    features: int,
    vector_distance: str = "euclidean",
    map_distance: str = "euclidean_toroid",
    neighborhood_kernel: str = "gaussian",
    use_cuda: bool = False,
    use_local_neighborhood: bool = False,
    use_fastmath: bool = True,
    save_progress: Optional[str] = None,
    low: float = 0.0,
    high: float = 1.0,
    seed: Optional[int] = None,
)

Main Class to create and train a Self-Organizing Map

Parameters:

  • rows (int) –

    Number of rows of neurons.

  • columns (int) –

    Number of columns of neurons.

  • features (int) –

    Numbers of features to the data / weights of each neuron.

  • vector_distance (str, default: 'euclidean' ) –

    Distance used in original data space, by default "euclidean". Possible values: "euclidean", "manhattan", "cosine"

  • map_distance (str, default: 'euclidean_toroid' ) –

    Distance used in map space, by default "euclidean_toroid". Possbile values: "euclidean", "manhattan", "euclidean_toroid", "manhattan_toroid"

  • neighborhood_kernel (str, default: 'gaussian' ) –

    Shape of the neighborhood kernel, by default "gaussian".

  • use_cuda (bool, default: False ) –

    If True, CUDA accelleration is used. Needs numba-cuda-mlir, installed via the cu12 or cu13 extra. By default False.

  • use_local_neighborhood (bool, default: False ) –

    Sets a hard neighborhood cutoff, by default False. Only used on CPU. Significantly increases performance at cost of numerical accuracy.

  • use_fastmath (bool, default: True ) –

    Slightly decrease numerical accuracy to increase performance, by default True.

  • save_progress (Optional[str], default: None ) –

    Saves codebook and U-Matrix to the given location if set, by default None. Usefull if long running computations crash / time out.

  • low (float, default: 0.0 ) –

    Lower bound for codebook initialization, by default 0.0.

  • high (float, default: 1.0 ) –

    Upper bound for codebook initialization, by default 1.0.

  • seed (Optional[int], default: None ) –

    Randomness seed for replicability, by default None.

Raises:

  • ValueError –

    If the map dimensions are less than 1x1.

  • ValueError –

    If the number of features is less than 2.

  • ImportError –

    If CUDA is requested but not available.

  • ValueError –

    If the vector distance norm is not one of the supported norms.

  • ValueError –

    If the map distance norm is not one of the supported norms.

  • ValueError –

    If the neighborhood kernel is not one of the supported kernels.

Methods:

  • predict –

    Return the positions of the BMU for a dataset

  • train –

    Train the SOM with the given data for a number of epochs.

  • train_batch –

    Manually train the SOM with a single batch of data.

Attributes:

  • u_graph (Graph) –

    The U-Distance graph for the current codebook.

  • umatrix (UMatrix) –

    The U-Matrix for the SOM, as a 3D array of shape

u_graph property

u_graph: Graph

The U-Distance graph for the current codebook.

Undirected, edge-weighted networkx graph whose nodes are grid positions (row, column) and whose edges connect toroidal grid neighbors, weighted by the high-dimensional distance between the corresponding codebook vectors, the same per-neighbor distances used to build the U-Matrix. The graph is cached, but bind it to a local name and reuse that for repeated chisom.analysis.u_distance (or future ranking) queries.

Returns:

  • Graph –

    The U-Distance graph for the current codebook.

umatrix property

umatrix: UMatrix

The U-Matrix for the SOM, as a 3D array of shape (n_layers, rows, columns).

Computed and cached after training. When save_progress is set, one layer is stored per epoch (training history); otherwise a single final layer is stored. Accessing this before training computes a single layer from the current (untrained) codebook on demand.

Returns:

  • NDArray[float16] –

    The U-Matrix, shape (n_layers, rows, columns).

predict

predict(
    data: NDArray | DataLoader,
) -> Tuple[NDArray[np.uint16], NDArray[np.float32]]

Return the positions of the BMU for a dataset

Parameters:

  • data (NDArray | DataLoader) –

    Dataset to find the BMUs for.

Returns:

  • NDArray[uint16] –

    The BMUs for the data.

  • NDArray[float32] –

    The Quantization Error

Raises:

  • TypeError –

    Error if the data format is not known

train

train(
    data: NDArray | DataLoader,
    epochs: int,
    alpha: float,
    batchsize: int = 1,
    shuffle: bool = True,
    sigma: Optional[int] = None,
    alpha_decay: str = "linear",
    sigma_decay: str = "exponential",
    alpha_end: float = 0.01,
    sigma_end: int = 1,
) -> None

Train the SOM with the given data for a number of epochs.

Runs the full training loop internally, including per-step decay of alpha and sigma and (if save_progress was set on initialization) periodic U-Matrix/codebook checkpointing once per epoch. Progress is reported via tqdm progress bars for both the epoch loop and the per-epoch batch loop.

Parameters:

  • data (NDArray | DataLoader) –

    The data to train the SOM with. If a DataLoader is used, it is iterated batch by batch as configured on the DataLoader itself. If a numpy array is used, it is split into batches of size batchsize.

  • epochs (int) –

    Number of epochs to train for.

  • alpha (float) –

    The initial learning rate. Decays to alpha_end over the course of training according to alpha_decay.

  • batchsize (int, default: 1 ) –

    Number of data points per batch. Only used when data is a numpy array (ignored for DataLoader input, which defines its own batching), by default 1.

  • shuffle (bool, default: True ) –

    If True, the order of batches is reshuffled at the start of each epoch. Only applies when data is a numpy array; DataLoader shuffling is controlled by the DataLoader itself. By default True.

  • sigma (Optional[int], default: None ) –

    The initial neighborhood radius. Must be greater than 0 if given. Defaults to None, in which case it is set to half the smaller of the map's row/column dimensions.

  • alpha_decay (str, default: 'linear' ) –

    The decay schedule for alpha, one of "linear" or "exponential", by default "linear".

  • sigma_decay (str, default: 'exponential' ) –

    The decay schedule for sigma, one of "linear" or "exponential", by default "exponential".

  • alpha_end (float, default: 0.01 ) –

    The value alpha decays towards by the end of training, by default 0.01.

  • sigma_end (int, default: 1 ) –

    The value sigma decays towards by the end of training, by default 1.

Raises:

  • ValueError –

    If sigma is given and is less than or equal to sigma_end.

  • ValueError –

    If alpha is given and is less than or equal to alpha_end.

  • ValueError –

    If alpha_decay or sigma_decay is not "linear" or "exponential".

train_batch

train_batch(
    data: NDArray, sigma: int, alpha: float
) -> None

Manually train the SOM with a single batch of data.

Provides fine-grained control over training by allowing the caller to set alpha and sigma explicitly for a single batch, bypassing the automatic per-epoch decay scheduling used by train. Useful for custom training loops or schedules not covered by train.

Parameters:

  • data (NDArray) –

    A single batch of vectors to train the SOM on.

  • sigma (int) –

    The neighborhood radius to use for this batch.

  • alpha (float) –

    The learning rate to use for this batch.

start_chisom_viewer

start_chisom_viewer(
    umatrix: Optional[NDArray] = None,
    bmu_coordinates: Optional[NDArray] = None,
    data: Optional[DatasetBase | DataFrame] = None,
    structure_info_column: Optional[str] = None,
    scaling_factor: int = 3,
)

Start the GUI interface

All arguments are optional; anything not passed here can be loaded from the viewer's File menu instead.

Parameters:

  • umatrix (Optional[NDArray], default: None ) –

    U-Matrix of the SOM.

  • bmu_coordinates (Optional[NDArray], default: None ) –

    Coordinates of the BMU to the data points.

  • data (Optional[DatasetBase | DataFrame], default: None ) –

    Additional data to the data points. Will be renderd in the tabel view and used for coloring of BMUs.

  • structure_info_column (Optional[str], default: None ) –

    With chemical dataset the column with this index supplies the SMILES to render the molecule, by default None.

  • scaling_factor (int, default: 3 ) –

    Will scale the U-Matrix by this factor ands interpolation for an anti-aliased view, by default 3.

chisom.utils

Utility tools for hyperparameter configuration for (emergent) SOMs

Functions:

  • decay_exponential –

    Calculate the exponential decay of a value towards a desire final value for use in the training process.

  • decay_linear –

    Calculate the linear decay of a value for use in the training process.

  • lattice_size –

    Returns a rectangular lettice for the given number of data points,

decay_exponential

decay_exponential(
    iteration: int,
    initial_value: int | float,
    end_value: Optional[int | float] = None,
    total_iterations: Optional[int] = None,
    decay: Optional[float] = None,
    *args,
    **kwargs,
) -> float

Calculate the exponential decay of a value towards a desire final value for use in the training process. Ether to use a fixed number of iterations or a decay factor. 'decay' takes precedence over 'total_iterations'.

Parameters:

  • iteration (int) –

    Current iteration.

  • initial_value (int | float) –

    Starting value of the decay.

  • end_value (Optional[int | float], default: None ) –

    Desired final value, when not using decay by default None.

  • total_iterations (Optional[int], default: None ) –

    Total number of desired iterations, by default None.

  • decay (Optional[float], default: None ) –

    Decay rate, by default None.

Returns:

  • float –

    Value at iteration iteration.

Raises:

  • ValueError –

    If neither total_iterations nor decay is provided, or if both end_value and decay are provided.

decay_linear

decay_linear(
    iteration: int,
    initial_value: int | float,
    total_iterations: Optional[int] = None,
    decay: Optional[float] = None,
    *args,
    **kwargs,
) -> float

Calculate the linear decay of a value for use in the training process. Ether to use a fixed number of iterations or a decay factor. 'decay' takes precedence over 'total_iterations'.

Parameters:

  • iteration (int) –

    Current iteration.

  • initial_value (int | float) –

    Starting value of the decay.

  • total_iterations (Optional[int], default: None ) –

    Total number of desired iterations, by default None.

  • decay (Optional[float], default: None ) –

    Decay rate, like m in y = -mx+b, by default None.

Returns:

  • float –

    Value at iteration iteration.

Raises:

  • ValueError –

    If neither total_iterations nor decay is provided.

lattice_size

lattice_size(
    dataset_size: int, factor=3
) -> Tuple[int, int]

Returns a rectangular lettice for the given number of data points, as recommendet by Ultsch et al. for ESOMs.

Parameters:

  • dataset_size (int) –

    Number of data points in the dataset.

  • factor –

    Ration of neurons to data points, by default 3.

Returns:

  • Tuple[int, int] –

    Number of (rows, columns)

chisom.analysis

Post-training analysis of a trained SOM: distance and (future) ranking queries over the SOM's U-Distance graph (see chisom.Som.u_graph).

Functions:

  • u_distance –

    Compute the u-distance(s) from source position(s) to a target position.

u_distance

u_distance(
    graph: Graph,
    source: Optional[
        PositionLike | Collection[PositionLike]
    ],
    target: PositionLike,
) -> dict[tuple[int, int], float]

Compute the u-distance(s) from source position(s) to a target position.

The u-distance is the shortest-path length between a source and target through a U-Distance graph (see chisom.Som.u_graph), where edges are weighted by the high-dimensional distance between neighboring codebook vectors. Build the graph once via chisom.Som.u_graph and reuse it across many u_distance calls (e.g. for many BMU pairs) rather than rebuilding it per call.

Parameters:

  • graph (Graph) –

    U-Distance graph, as returned by chisom.Som.u_graph.

  • source (Optional[PositionLike | Collection[PositionLike]]) –

    Grid position(s) to compute the distance from: - None: every node in graph. - A single position (tuple/array of 2 ints): that one node. - A collection of positions: each given source node.

  • target (PositionLike) –

    Grid position (row, column) to compute the distance to. Same accepted forms as a single source position.

Returns:

  • dict[tuple[int, int], float] –

    Mapping from each queried source node (as an int (row, column) tuple) to its u-distance to target. Has one entry per node in graph when source is None, a single entry for a single source position, or one entry per given source for a collection thereof.

Datasets are loaded by suffix. chisom.io.load_dataset (and the viewer's File → Load data dialog) accept .h5/.hdf5 for HDF5 stores, .csv for comma-separated text, .tsv/.txt for tab-separated text, and .parquet/.pq for Parquet. The full tuple is available as chisom.io.loading.DATASET_SUFFIXES.

chisom.io

Classes and Functions for large on-disk data stores specific to cheminformatics

HDF5Creator

HDF5Creator(
    fingerprint_generator_factory: DataloaderFingerprintGeneratorFactory,
    file_extensions: list[str] = [".txt", ".smi", ".csv"],
    num_processes: int = mp.cpu_count() - 2,
    chunk_size: int = 1000,
    queue_size: int = 500,
)

Bases: StoreCreator

HDF5 Store Create for creating HDF5 file of cheminformatic datasets from file hierarchies.

Parameters:

  • fingerprint_generator_factory (DataloaderFingerprintGeneratorFactory) –

    Factory to use, depends on the input file type and content.

  • file_extensions (list[str], default: ['.txt', '.smi', '.csv'] ) –

    File extentions to consider.

  • num_processes (int, default: cpu_count() - 2 ) –

    Number of processes used.

  • chunk_size (int, default: 1000 ) –

    Size of chunks send to processes. Should not need configuration.

  • queue_size (int, default: 500 ) –

    Size of queue of each process,. Should not need configuration.

create

create(
    file_hierarchy: FileList,
    out_path: str,
    leaf_map: LeafMap,
    skip_lines: int = 0,
    sep: str = "\t",
) -> None

Run creation of HDF5 storage file

Parameters:

  • file_hierarchy (FileList) –

    Dictionary of files to parse, see How-To Guides.

  • out_path (str) –

    Output path for the HDF5 files.

  • leaf_map (LeafMap) –

    Dictionary of file structure and datatypes, see How-To Guides.

  • skip_lines (int, default: 0 ) –

    Number of lines to skip at the beginning of file, e.g. for headers, by default 0.

  • sep (str, default: '\t' ) –

    Column seperator, by default "\t".

CSVStyleFactory

CSVStyleFactory(
    mol_generator: MolGenerator,
    fpStart: int,
    fpSize: int,
    dtype: type = np.float32,
)

Bases: DataloaderFingerprintGeneratorFactory

A generator that creates fingerprints from a row of data, not from a Mol object. With a similar interface to the RDKit fingerprint generator, but using a slice of the row instead of a Mol object.

Parameters:

  • mol_generator (MolGenerator) –

    The RDKit MolGenerator to use to read Molecules in the input file e.g. MolFromSmiles

  • fpStart (int) –

    Numerical column index where the fingerprints starts.

  • fpSize (int) –

    Length in number of columns of the CSV file of the fingerprint.

  • dtype (type, default: float32 ) –

    Datatype of the fingperint elements

rdStyleFactory

rdStyleFactory(
    mol_generator: MolGenerator,
    fingerprint_generator: rdFingerprintGenerator,
    generator_kwargs={},
    count_fingerprint=False,
)

Bases: DataloaderFingerprintGeneratorFactory

This class is designed to provide a unified interface for generating molecular fingerprints for use in DataLoader creation. It can combine different types of input (SMILES, InChI, ...) and fingerprint generator (e.g., Morgan, RDKit, Feature Morgan ... ). Also returns a new instance of the real generator on each call to get_generator, so that the generator can be used in a multiprocessing context without issues.

Parameters:

  • mol_generator (MolGenerator) –

    The RDkit MolGenerator to use to read Molecules in the input file. e.g. MolFromSmiles

  • fingerprint_generator (rdFingerprintGenerator) –

    The rdFingerprintGenerator to use.

  • generator_kwargs (dict, default: {} ) –

    Keyword argument to pass to the rdFingerprintGenerator.

  • count_fingerprint (bool, default: False ) –

    If count fingerprints should be created where applicable.

HDF5Dataset

HDF5Dataset(
    filepath: str, group_subset: Optional[List[str]] = None
)

Bases: DatasetBase

Loads HDF5 Datasets for large cheminformatic datasets created with the supplied script. Adheres to the PyTorch Dataset interface for the use with the PyTorch DataLoader for milisecond on-disc access data access.

Parameters:

  • filepath (str) –

    Path to HDF5 file

  • group_subset (Optional[List[str]], default: None ) –

    Subsets to use, by default None

dataset_column_names

dataset_column_names(
    data: Union[DatasetBase, DataFrame],
) -> List[str]

Return the column names of either dataset flavour.

Parameters:

  • data (Union[DatasetBase, DataFrame]) –

    A dataset as returned by load_dataset.

Returns:

  • List[str] –

    The column names.

inspect_hdf5_groups

inspect_hdf5_groups(
    filepath: Union[str, Path],
) -> List[str]

Read the group names of an HDF5 store without loading the dataset.

Parameters:

  • filepath (Union[str, Path]) –

    Path to the HDF5 store.

Returns:

  • List[str] –

    The names of the groups at the root of the store, sorted like HDF5Dataset sorts them.

load_bmu_coordinates

load_bmu_coordinates(filepath: Union[str, Path]) -> NDArray

Load BMU coordinates from a .npy file.

Parameters:

  • filepath (Union[str, Path]) –

    Path to a NumPy array of shape (n_datapoints, 2) holding the row/column coordinate of the best matching unit of every datapoint.

Returns:

  • NDArray –

    The BMU coordinates.

Raises:

  • ValueError –

    Raised if the array does not have the expected shape or dtype.

load_dataset

load_dataset(
    filepath: Union[str, Path],
    group_subset: Optional[List[str]] = None,
) -> Union[DatasetBase, DataFrame]

Load a dataset of datapoint properties, dispatching on the file extension.

Parameters:

  • filepath (Union[str, Path]) –

    Path to the dataset. Supported are HDF5 stores created with HDF5Creator (.h5, .hdf5), delimited text (.csv, .tsv, .txt) and Parquet (.parquet, .pq).

  • group_subset (Optional[List[str]], default: None ) –

    HDF5 only: the groups to include, by default all of them.

Returns:

  • Union[DatasetBase, DataFrame] –

    An HDF5Dataset for HDF5 stores, a DataFrame otherwise.

Raises:

  • ValueError –

    Raised if the file extension is not supported.

load_umatrix

load_umatrix(filepath: Union[str, Path]) -> NDArray

Load a U-matrix from a .npy file.

Parameters:

  • filepath (Union[str, Path]) –

    Path to a NumPy array holding a 2D (single layer) or 3D (layered) U-matrix.

Returns:

  • NDArray –

    The U-matrix, always as a 3D array with the layer as first axis.

Raises:

  • ValueError –

    Raised if the file does not contain a 2D or 3D array.

plot_som

plot_som(
    umatrix: UMatrix,
    bmu_coordinates: Optional[NDArray[uint16]] = None,
    data: Optional[Union[DatasetBase, DataFrame]] = None,
    color_by: Optional[str] = None,
    *,
    categorical: Optional[bool] = None,
    category_colors: Optional[Mapping[Any, Any]] = None,
    cmap: Union[str, Colormap] = "viridis",
    umatrix_cmap: Optional[Union[str, Colormap]] = None,
    alpha_scheme: Optional[str] = "excess_absolute",
    layer: int = -1,
    scaling_factor: int = 3,
    marker_size: Optional[float] = None,
    marker_cell_fraction: float = 0.7,
    legend_ncol: Optional[int] = None,
    legend_label_maxlen: int = 24,
    chrome_scale: float = 1.0,
    figsize: Optional[tuple[float, float]] = None,
    cell_size_in: float = DEFAULT_CELL_SIZE_IN,
    dpi: int = 150,
    ax: Optional[Axes] = None,
    save_as: Optional[Union[str, Path]] = None,
) -> Figure

Plot a trained SOM as a static matplotlib figure.

Draws the interpolated U-matrix as background image and, if BMU coordinates are given, the BMUs as scatter markers on top — optionally colored by a property of the underlying data. Reproduces the default look of the interactive viewer, but works headless (e.g. with the Agg backend) and can be written directly to file.

Sizing of the map image, the BMU markers, and the colorbar(s)/legend is coupled by default: the figure is sized from the SOM's actual grid shape (cell_size_in), the axes box is aspect-locked to that same grid shape (rows / columns), and marker size is derived from how large a single grid cell actually renders once that layout is resolved. This keeps the map, its markers, and its chrome proportionate to each other regardless of how big or how non-square a given SOM's lattice is. Pass explicit figsize/marker_size to opt back into fixed, grid-independent sizing.

Parameters:

  • umatrix (UMatrix) –

    U-matrix as returned by the Som.umatrix property. Either 2D (rows, columns) or 3D (layers, rows, columns).

  • bmu_coordinates (NDArray[uint16], default: None ) –

    (N, 2) array of (row, column) BMU coordinates as returned by Som.predict(). If None, only the U-matrix is drawn.

  • data (DatasetBase or DataFrame, default: None ) –

    Data source holding the property values referenced by color_by. Must have one entry per row in bmu_coordinates.

  • color_by (str, default: None ) –

    Name of the column in data to color the BMUs by. If None, BMUs are drawn in plain black.

  • categorical (bool, default: None ) –

    Force categorical (True) or continuous (False) coloring. By default this is taken from the column properties of data (for a plain DataFrame the heuristic of the viewer is used: at most 10 unique values within the first 100 rows counts as categorical).

  • category_colors (Mapping, default: None ) –

    Mapping of category value to matplotlib color, used for categorical coloring. By default colors are assigned from the tab10/tab20 palettes in sorted category order.

  • cmap (str or Colormap, default: 'viridis' ) –

    Matplotlib colormap for continuous property coloring.

  • umatrix_cmap (str or Colormap, default: None ) –

    Matplotlib colormap for the U-matrix. Defaults to the Earth colormap of the viewer.

  • alpha_scheme ((gini, excess_absolute, excess_relative), default: "gini" ) –

    Weighting scheme used to encode the dominance of the primary category as marker opacity (see RatioWeighting). If None, markers are fully opaque. Ignored for continuous coloring.

  • layer (int, default: -1 ) –

    Layer to display for a 3D U-matrix.

  • scaling_factor (int, default: 3 ) –

    Upscaling factor for the U-matrix interpolation.

  • marker_size (float, default: None ) –

    BMU marker diameter in points. By default (None), it is derived from the rendered size of one SOM grid cell (see marker_cell_fraction) so markers stay proportionate to the map regardless of grid shape or figure size. Pass a float to pin an exact, grid-independent size instead.

  • marker_cell_fraction (float, default: 0.7 ) –

    When marker_size is auto-derived, the fraction of one rendered grid cell's size (the smaller of its width/height in points) that the marker diameter should occupy. Ignored if marker_size is set.

  • legend_ncol (int, default: None ) –

    Number of columns for the categorical-coloring legend. By default (None), chosen automatically so the legend wraps into additional columns rather than growing arbitrarily tall for large category counts (see _auto_legend_ncol).

  • legend_label_maxlen (int, default: 24 ) –

    Maximum rendered length of a categorical legend label before it is elided with "…", so a few very long category names can't blow up the legend's (and therefore the map's) reserved width.

  • chrome_scale (float, default: 1.0 ) –

    Single multiplier applied together to the auto-derived marker size, the colorbar's fraction/pad, and the legend/colorbar font size, and to the chrome width reserved in an auto-derived figsize. Use this to scale all non-map chrome up or down in one step (e.g. for a much larger or smaller figsize) instead of tuning each piece of chrome separately. Does not affect an explicitly-passed marker_size or figsize.

  • figsize (tuple of float, default: None ) –

    Figure size in inches, passed to matplotlib. By default (None), derived from the SOM's grid shape via cell_size_in instead of matplotlib's own grid-independent default.

  • cell_size_in (float, default: DEFAULT_CELL_SIZE_IN ) –

    Inches per SOM grid cell, used to derive figsize when figsize is not given. Ignored if figsize is set explicitly.

  • dpi (int, default: 150 ) –

    Figure resolution, passed to matplotlib.

  • ax (Axes, default: None ) –

    Existing axes to draw into. If given, figsize, dpi, and cell_size_in are ignored and the legend is placed inside the axes.

  • save_as (str or Path, default: None ) –

    If given, the figure is saved to this path; the format is inferred from the suffix (e.g. .pdf).

Returns:

  • Figure –

    The matplotlib figure, for further customization.