Clustering is an invaluable cheminformatics technique for subdividing a typically large compound collection into small groups of similar compounds. One of the advantages is that once clustered you can store the cluster identifiers and then refer to them later this is particularly valuable when dealing with very large datasets. This often used in the analysis of high-throughput screening results, or the analysis of virtual screening or docking studies. A while back I looked at the options for clustering large datasets.

I thought it might be interesting to compare the performance of various clustering approaches on Apple Silicon (Apple M2 Ultra Mac with 192 GB of RAM.).

These four Jupyter notebooks cluster a set of molecules by structural similarity, using RDKit Morgan fingerprints (radius 3, 2048 bits) and Tanimoto similarity.

NotebookMethodUse it for
BitBirchclustering.ipynbBitBIRCHAny size, from hundreds to hundreds of thousands of molecules
Butinaclustering.ipynbRDKit Taylor–Butina (rdkit.ML.Cluster.Butina)Small to medium sets; memory grows with the square of the set size
MLXButinaclustering.ipynbTaylor–Butina on the Apple-silicon GPU (mlxmolkit)The same clusters as RDKit Butina, about 6× faster and in about 20× less memory at 150k molecules. Needs macOS on Apple silicon.
KMeansclustering.ipynbk-means (scikit-learn)Splitting a set into a chosen number of groups, e.g. for diverse subsets or train/test splits. It suggests a number of clusters and lets you change it.

All four notebooks have the same inputs, outputs and layout, so you can run a dataset through any of them and compare the results. BitBIRCH and the two Butina notebooks group molecules by a similarity threshold. k-means instead uses a number of clusters you choose. All notebook use the same conda enviroment. Passing a square matrix to Butina.ClusterData needs a recent RDKit, so environment.yml requires RDKit 2026.03 or newer.

Choosing a threshold

Diverse sets, such as approved drugs or random screening compounds, give mostly single-molecule clusters at 0.65. Around 0.35–0.45 gives useful families. Some reference points:

Data setThresholdClusters (BitBIRCH)Largest
789 approved drugs0.657693
789 approved drugs0.356329
149k random ZINC compounds0.65104,93530

Performance and memory

Timings are from an Apple M2 Ultra Mac with 192 GB of RAM.

All four notebooks were run on the same two sets: 149,396 random ZINC compounds and the 4,629-molecule ERK set. The clustering time is the clustering step alone. The notebook time includes loading the file (about 50 s for 150k), fingerprints and export. “Agreement” is the adjusted Rand index against RDKit Butina at the same threshold (1 = identical partition, 0 = no better than chance).

150k molecules:

MethodSettingClusteringNotebookPeak RAMClustersSingle-moleculeLargestAgreement with RDKit Butina
RDKit Butinasimilarity ≥ 0.65139 s191 s97 GB89,46454,53540reference
MLX Butina (GPU)similarity ≥ 0.6523 s76 s4.6 GB89,23354,393400.94
BitBIRCHsimilarity ≥ 0.6523 s76 s12 GB104,93575,886300.54
RDKit Butinasimilarity ≥ 0.35139 s190 s97 GB23,0647,525346reference
MLX Butina (GPU)similarity ≥ 0.3523 s74 s4.8 GB21,3627,4573460.76
BitBIRCHsimilarity ≥ 0.3521 s74 s10 GB68,79541,7861270.12
k-meansk = 103 (suggested)39 s, plus a 181 s k scan273 s6.6 GB103518,775—
k-meansk = 250109 s163 s6.4 GB2502344,183—

An optional t-SNE map adds about 4 minutes at 150k, and about 20 s for the ERK set, in every notebook.

ERK set (4,629 molecules): every method clusters in 0.2–0.6 s, apart from k-means (2 s at k = 41, 14 s at k = 250, plus 23 s for the k scan). Peak RAM is under 1.2 GB. The agreement with RDKit Butina is 0.95 for MLX Butina at both thresholds, and 0.60 (at 0.65) and 0.34 (at 0.35) for BitBIRCH.

What this shows:

  • MLX Butina gives Butina results at BitBIRCH cost. At 150k it is about 6× faster than RDKit Butina and uses about 20× less memory, because it keeps only the neighbour pairs instead of the full distance matrix. Its GPU memory peak was only 0.16 GB.
  • MLX and RDKit Butina differ only in the greedy step. The neighbour lists are identical: every pair above the threshold matches RDKit’s Tanimoto exactly. mlxmolkit recounts neighbours after each cluster (RDKit’s reordering=True) and breaks ties the other way. RDKit’s notebook defaults to reordering=False. The difference is small at 0.65 (0.94–0.95 agreement, the same largest clusters). It grows at 0.35 on 150k (0.76, 21,362 against 23,064 clusters), where recounting lets later clusters absorb more molecules.
  • BitBIRCH is fast but partitions differently. It makes more, smaller clusters than Butina at the same threshold, and agrees less with it at low thresholds.
  • k-means isn’t suited to diverse sets. MiniBatchKMeans on 150k diverse molecules produces one huge catch-all cluster (19k molecules at k = 103, 44k at k = 250) alongside many small ones. Use it to split a set into a fixed number of groups, not to find families.

MLX Butina check. Before clustering, the MLX notebook checks that a Metal GPU is available. It estimates memory from the neighbour density in a sample, which matters most at low thresholds on sets of close analogues. It also times two trial runs to give a rough, usually high, estimate of the run time.

Butina memory check. Butina needs every pairwise distance, and RDKit holds them in a full n × n matrix. The notebook builds this matrix as float32 in blocks, which uses about 5× less memory than RDKit’s usual flat distance list. Before clustering starts, a check runs that:

  • estimates the peak memory, including the matrix and neighbour lists
  • compares it with the RAM available at that moment
  • prints the largest set size that would fit, and an estimate of the run time
  • stops with an error if the estimate is more than MEMORY_FRACTION of the available RAM

Related Posts