Auto3D 3.0.0
First release published to PyPI since 2.3.1. The v3.0.0 and v3.5.0 git tags that existed previously were never published to any index; if you installed with pip or conda, every change below applies to you.
Full details in CHANGELOG.md; upgrade path in the migration guide.
Results that change
- Thermochemistry now uses most-abundant-isotope masses, matching Gaussian and ORCA. Every reported
H_hartree,S_hartree_per_KandG_hartreechanges. Mass enters the moments of inertia, the mass-weighted Hessian (so every frequency, ZPE and Svib) and the translational term. Effect grows with heavy-halogen content. Cached or published thermochemistry from an earlier version is no longer bit-comparable — regenerate. E_totis Hartree in every file Auto3D writes. It previously meant eV or Hartree depending on which component wrote the file, while consumers hard-coded eV.opt_stepsbelow 10 is refused at construction. It was returning an unconverged structure labelled as optimized: the optimizer only tests convergence onistep % 10 == 0and guards its statistics withn >= 10.
Fixed
- A custom NNP returning float64 forces crashed the FIRE loop — but only when two or more molecules reduced in the same step, so it was never reliably reproducible.
smiles2molsbuilttorch.device("cuda:N")with no bounds check, so an out-of-range GPU index producedcuda:99instead of an error.reorder_sdfstaged through a predictable temp name and dropped the target's file mode, so a0600input came back0644— a permission loosening.select_tautomerstruncated its derived output path with no gate.- Blank lines in a
.smifile raised a bareIndexErroror tuple-unpackValueError. - A malformed config file exited 1 through
auto3d <config>.yamlbut 2 throughauto3d run -c, so a script gating on the exit code got different answers from the two entry points. - Chunk sizing initialized a full CUDA context in the orchestrator process before workers were spawned.
- The OOM retry ran
empty_cache()while the exception was still live — so it could only release already-free blocks — and then reused the original unshrunk batch size for later slices.
Performance
-
import Auto3Dis 44x cheaper: 1.35 s → 0.031 s, 1175 modules → 154. It no longer imports torch or RDKit. Three eager optional-dependency probes were defeating the lazy-import mechanism that exists to prevent exactly this. -
The optimizer's host↔device synchronizations drop from 18 per step to 2, and ANI2xt's from 22 per forward to 2. Proven bit-identical against a reference implementation of the old loop across 17 scenarios. The ANI2xt element loop now compiles to one
torch.compilegraph instead of zero (it was skipping the whole frame on a data-dependent branch).No wall-clock speedup is claimed. Every count above is CI-enforced on CPU; durations require a GPU and have not been measured.
benchmarks/run_perf_ab.shis provided for that. Documentation that previously claimed a ~1.25xtorch.compilegain has been corrected — that path was compiling to zero graphs, so the figure could not have come from it.
Breaking API changes
from Auto3D.utils import <name>no longer resolves for 41 names;utils/chemistry.pyandutils/file_ops.pyare split up. The CHANGELOG carries the full old→new mapping.optimizingtakes a built adapter, keyword-only, and no longer resolves a model name itself. Construct the adapter inside the worker process.- One model contract:
ModelAdapterlives inAuto3D.models.contract.CustomNNP— the public custom-NNP contract — is unchanged. check_valid_configurationtakes anAuto3DOptionsinstead of ten keyword arguments.filter_uniqueis deleted; the two conformer filters are one. Filters return aFilterResultthat records why conformers were dropped, so "No structure converged" is no longer reported when the cause was stereochemistry.smiles2smi,encode_ids,decode_idsandselect_tautomerstake a keyword-onlyoverwrite.Auto3D.cliandAuto3D.modelsno longer re-export names;cli/__init__.pywas additionally shadowing its ownappandconsolesubmodules.
Tests
1273 → 1646, with no expected failures. Four tests that could not fail were found and fixed, one of which had been concealing the exit-code divergence above by re-implementing inline the very function it claimed to exercise.