Building QupKake on Apple Silicon (macOS arm64)

QupKake predicts micro-pKa values with GFN2-xTB plus graph neural networks. Its README targets Linux. On an Apple Silicon Mac the obvious routes all fail:

  • The bundled xtb-641 binary is Linux x86-64 and can’t run on macOS.
  • conda-forge xtb is rejected by QupKake on import (“Conda version of xTB is currently not supported”). The conda-forge build also fails at runtime with Library not loaded: @rpath/libgfortran.5.dylib.
  • A source build of xtb 6.4.1 with current compilers builds and runs, but at the default -O3 it is silently miscompiled: geometry optimizations climb in energy and never converge (details under Pitfalls).

This guide gives a setup that works, tested on macOS 26.6 (Mac Studio, M-series, 24 cores) in October 2026.

ComponentVersion used
Python3.9 (conda-forge)
PyTorch / torchvision2.1.2 / 0.16.1 (conda-forge, CPU)
pytorch-lightning / torchmetrics2.1.4 / 1.8.2
torch_geometric2.4.0 (pip)
RDKit2025.03.5
xtb6.4.1 from source, gfortran 16.2 (conda-forge), OpenBLAS 0.3.34, meson 1.8

Throughout this guide, QK is the directory the QupKake repo is cloned into:

export QK=$HOME/Projects/QupKake

1. Clone QupKake

git clone https://github.com/Shualdon/QupKake.git "$QK"

2. Create the conda environment

Use conda-forge only. The pytorch channel named in the repo’s environment.yaml no longer publishes builds. The same environment also supplies the Fortran toolchain for xtb, so the binary links against the env’s own libgfortran and openblas.

conda create -n qupkake -c conda-forge --override-channels -y \
  python=3.9 pip "pytorch=2.1" torchvision torchmetrics "pytorch-lightning=2.1" \
  pandas numpy scipy rdkit matplotlib jupyterlab ipykernel \
  meson ninja pkg-config gfortran openblas liblapack c-compiler
conda activate qupkake

pip install torch_geometric==2.4
pip install "$QK"

Torch is pinned to 2.1, as in the repo’s environment.yaml, to match the shipped model checkpoints.

3. Build xtb 6.4.1 from source

3a. Download and verify

mkdir -p "$QK/xtb-src" && cd "$QK/xtb-src"
curl -LO https://github.com/grimme-lab/xtb/releases/download/v6.4.1/xtb-6.4.1.tar.xz
curl -LO https://github.com/grimme-lab/xtb/releases/download/v6.4.1/xtb-6.4.1.tar.xz.sha256
shasum -a 256 xtb-6.4.1.tar.xz; cat xtb-6.4.1.tar.xz.sha256   # must match:
# 15cfd8743dd963c54e0fee12da5ee34a08dde78625cf8a0a21eee1eaa652beae
tar xf xtb-6.4.1.tar.xz && cd xtb-6.4.1

3b. Patch one format string

gfortran 16 rejects a malformed format descriptor at runtime: Fortran runtime error: Missing comma between descriptors. xtb hits it during every --opt run. The patch adds the missing commas:

sed -i '' '624s/(1x,"("f7.2"%)")/(1x,"(",f7.2,"%)")/' src/optimizer.f90
sed -n 624p src/optimizer.f90
# expected:       write(env%unit,'(1x,"(",f7.2,"%)")')         (depred-echng)/echng*100

3c. Configure and build at -O2 (not the default -O3)

FC=gfortran CC=cc meson setup build-o2 \
  --prefix="$QK/xtb-6.4.1" \
  --buildtype=debugoptimized \
  -Dla_backend=openblas -Dopenmp=true
ninja -C build-o2 install
  • debugoptimized means -O2 -g. Do not use --buildtype=release. It gives -O3, which gfortran 16 miscompiles.
  • la_backend=openblas: the default is mkl-static, which doesn’t exist on Apple Silicon.
  • meson prints WARNING: FC and CC are not from the same vendor. That’s expected and harmless.

The installed binary finds libgfortran, libopenblas and libomp through rpaths into the conda env. To confirm:

otool -l "$QK/xtb-6.4.1/bin/xtb" | grep -A2 LC_RPATH   # should list .../envs/qupkake/lib

3d. Verify with a real optimization

A --version check or a single-point energy is not enough: the miscompiled -O3 binary passes both. Run a full geometry optimization, in gas phase and in ALPB water (the settings QupKake uses):

cd "$(mktemp -d)"
python -c "
from rdkit import Chem; from rdkit.Chem import AllChem
m = Chem.AddHs(Chem.MolFromSmiles('CC(=O)Oc1ccccc1C(=O)O')); AllChem.EmbedMolecule(m, randomSeed=42)
AllChem.MMFFOptimizeMolecule(m); Chem.MolToMolFile(m, 'aspirin.mol')"
export OMP_STACKSIZE=4G
"$QK/xtb-6.4.1/bin/xtb" aspirin.mol --opt > gas.txt 2>&1
"$QK/xtb-6.4.1/bin/xtb" aspirin.mol --opt --alpb water --lmo -P 4 > water.txt 2>&1
grep -E 'OPTIMIZATION CONVERGED|FAILED TO CONVERGE|TOTAL ENERGY|abnormal' gas.txt water.txt

Expected result: both runs print GEOMETRY OPTIMIZATION CONVERGED after about 20–25 iterations, with total energies of about −39.632 Eh (gas) and −39.652 Eh (water). A broken build gives FAILED TO CONVERGE or abnormal termination instead, and its energy rises at every optimization step.

Shell gotcha when scripting tests: in zsh (the macOS default shell), an unquoted $args holding "--opt --alpb water" does not word-split. xtb receives one garbage argument, quietly ignores it and runs a gas-phase single point. Use bash, or ${=args} in zsh.

4. Point QupKake at xtb

QupKake reads XTBPATH once, at import time (qupkake/__init__.py), and it must be the path to the executable. Set it, along with OpenMP settings, in conda activation hooks so conda activate qupkake sets everything:

mkdir -p "$CONDA_PREFIX/etc/conda/activate.d" "$CONDA_PREFIX/etc/conda/deactivate.d"

cat > "$CONDA_PREFIX/etc/conda/activate.d/xtb.sh" <<EOF
export _QUPKAKE_OLD_OMP_STACKSIZE="\${OMP_STACKSIZE:-}"
export _QUPKAKE_OLD_OMP_NUM_THREADS="\${OMP_NUM_THREADS:-}"
export XTBPATH=$QK/xtb-6.4.1/bin/xtb
export XTBHOME=$QK/xtb-6.4.1/share/xtb
export OMP_STACKSIZE=4G
export OMP_NUM_THREADS=1
EOF

cat > "$CONDA_PREFIX/etc/conda/deactivate.d/xtb.sh" <<'EOF'
unset XTBPATH XTBHOME
if [ -n "$_QUPKAKE_OLD_OMP_STACKSIZE" ]; then export OMP_STACKSIZE="$_QUPKAKE_OLD_OMP_STACKSIZE"; else unset OMP_STACKSIZE; fi
if [ -n "$_QUPKAKE_OLD_OMP_NUM_THREADS" ]; then export OMP_NUM_THREADS="$_QUPKAKE_OLD_OMP_NUM_THREADS"; else unset OMP_NUM_THREADS; fi
unset _QUPKAKE_OLD_OMP_STACKSIZE _QUPKAKE_OLD_OMP_NUM_THREADS
EOF

conda deactivate && conda activate qupkake && echo "$XTBPATH"

xtb 6.4.1 has its GFN2 parameters compiled in, so XTBHOME is optional. It’s set anyway for completeness.

5. Jupyter kernel

Jupyter kernels don’t run conda activation hooks, so put the variables in the kernel spec. The kernel’s PATH must also include conda: on import QupKake runs conda list xtb, and if conda is missing the import crashes with a confusing NameError: name 'XTB_LOCATION' is not defined.

python -m ipykernel install --user --name qupkake --display-name "Python (qupkake)"
python - <<EOF
import json, os
p = os.path.expanduser("~/Library/Jupyter/kernels/qupkake/kernel.json")
k = json.load(open(p))
k["env"] = {
    "XTBPATH": "$QK/xtb-6.4.1/bin/xtb",
    "XTBHOME": "$QK/xtb-6.4.1/share/xtb",
    "OMP_STACKSIZE": "4G",
    "OMP_NUM_THREADS": "1",
    "PATH": "$CONDA_PREFIX/bin:$(dirname "$CONDA_EXE")/../condabin:/usr/bin:/bin:/usr/sbin:/sbin",
}
json.dump(k, open(p, "w"), indent=1)
EOF

A worked example notebook is in QupKake_usage.ipynb. It covers a single SMILES, a CSV batch, tautomers, site drawings and calling QupKake from Python.

6. Verify QupKake end to end

cd "$(mktemp -d)"
python -c "import qupkake; print(qupkake.XTB_LOCATION)"     # source-built xtb path
qupkake smiles "CC(=O)Oc1ccccc1C(=O)O" -n aspirin -r test
grep -A1 "<pka" test/output/qupkake_output.sdf

Expected result: it runs in about 6 s and predicts an acidic pKa of about 3.58 for aspirin’s carboxylic acid (experimental ≈ 3.5), plus a basic site at about 1.5.

Two harmless messages appear on every import:

  • CondaValueError: No packages match 'xtb' — QupKake checking for a conda xtb, which is deliberately absent.
  • GPU available but not used — QupKake runs on CPU.

Pitfalls

-O3 miscompilation (the main trap)

With gfortran 16 at -O3, xtb 6.4.1 runs, prints the right version and gives correct single-point energies. In geometry optimizations, though, the energy rises at every step until the SCF fails. QupKake reports this as:

RuntimeError: Problem running xtb on molecule
...
TypeError: 'NoneType' object is not iterable      (in predict_sites)

and the input SDF in <root>/raw/ is truncated to 0 bytes. -O2 and -O0 give identical, correct results. The cause wasn’t investigated further: a gfortran 16 optimizer bug or undefined behaviour in the 2021 xtb code.

-mp N oversubscribes the CPU

QupKake passes its worker count to every xtb call as -P N. -mp 16 therefore launches 16 xtb jobs with 16 threads each. For batch runs, point XTBPATH at a small wrapper that forces one thread per xtb process:

#!/bin/bash
# xtb1: force "-P 1" whatever QupKake asks for
args=(); skip=0
for a in "$@"; do
  if [ $skip = 1 ]; then args+=("1"); skip=0; continue; fi
  [ "$a" = "-P" ] && skip=1
  args+=("$a")
done
exec "$QK/xtb-6.4.1/bin/xtb" "${args[@]}"

Use it with XTBPATH=/path/to/xtb1 qupkake file mols.csv -mp 22 .... Without -mp, QupKake still passes -P <all cores>, so the wrapper helps there too.

Input structures

  • Give 2D SDFs as SMILES (CSV), not SDF. For SDF input QupKake passes the coordinates to xtb unchanged, with no hydrogens added. For CSV input it adds hydrogens and embeds in 3D first.
  • One unembeddable molecule kills the whole file. RDKit’s Bad Conformer Id aborts the run. Pre-check with qupkake.cli.embed_molecule and remove failures.
  • Output pKa values are sometimes the string tensor(4.21). Strip that before converting to float.
  • Tautomer and charge state matter. QupKake predicts micro-pKa of the structure as drawn, so a different tautomer gives different values. On a 4.6k-compound set:
    • RDKit’s canonical tautomer (rdMolStandardize.TautomerEnumerator.Canonicalize) changed the most acidic pKa by >1 unit for 18% of the 1,900 molecules it altered.
    • The registered structures agreed better with QupKake’s own xtb-based tautomer search (-t) than the RDKit canonical ones did.

-t (tautomer search)

  • -t runs one xtb optimization per RDKit-enumerated tautomer: a median of about 12 per molecule, up to 700+ for some. At about 4 CPU-seconds each, budget accordingly.
  • If xtb fails on any tautomer, QupKake aborts the molecule (KeyError: 'totalenergy'). A drop-in launcher that skips failed tautomers is in runs/ForPKA_taut/qupkake_t.py (use python qupkake_t.py file ... -t in place of qupkake file ... -t).
  • The GFN2-xTB/ALPB energy ranking sometimes favours implausible imidic-acid (C(OH)=N) and enol tautomers. Check its picks by eye.

Batch runs

For thousands of molecules, run many small single-worker QupKake jobs through xargs -P <cores>, not one big -mp run. QupKake splits a file into equal-count blocks per worker, so a single slow molecule stalls the whole run. Example scripts are in runs/ (ForPKA*/run_all.sh, run_job.sh, combine.py). Throughput on 22 cores is about 4.5 h per 4,600 drug-like molecules without -t.


Rebuilding xtb

conda activate qupkake
cd "$QK/xtb-src/xtb-6.4.1"
ninja -C build-o2 install

Then re-run the check in 3d.

I’ve uploaded details to GitHub https://github.com/drc007/qupkake-apple-silicon/tree/main/qupkake-apple-silicon

Citation

Abarbanel, O. D.; Hutchison, G. R. QupKake: Integrating Machine Learning and Quantum Chemistry for micro-pKa Predictions. J. Chem. Theory Comput. 2024. doi:10.1021/acs.jctc.4c00328

xtb: Bannwarth, C.; Ehlert, S.; Grimme, S. J. Chem. Theory Comput. 2019, 15, 1652. github.com/grimme-lab/xtb

Related Posts