-
Notifications
You must be signed in to change notification settings - Fork 6
ESM3Di CPU GPU benchmark
Report version 1.0.1 (release-status clarification; measurements unchanged). Measured on audrey1 on 2026-09-10 (JST), Slurm job 25616.
For the two proteins below (1378 residues total), median ESM3Di-35M inference time was 0.901931 seconds on CPU and 0.015422 seconds on CUDA: a 58.48× CPU/GPU elapsed-time ratio. Predictions were identical across CPU and CUDA and across all measured repetitions.
At measurement time, this used an unreleased CSUBST development worktree, based on commit
286ac2b,
with local 3Di predictor changes. It is not a benchmark of the published base
commit alone. The tested predictor source hashes below also match
CSUBST 1.15.0, which has
since been released on GitHub. Distribution availability is tracked in the
installation guide.
All times are seconds for the entire batch of two sequences.
| Device | Run 1 | Run 2 | Run 3 | Median |
|---|---|---|---|---|
| CPU (4 threads) | 0.901879553 | 0.905969813 | 0.901930879 | 0.901930879 |
| CUDA (RTX 6000 Ada) | 0.022375485 | 0.015422144 | 0.015390404 | 0.015422144 |
The ratio is CPU median / CUDA median, not a ratio of model-loading times. Only three repetitions were measured. CUDA values ranged from 0.015390 to 0.022375 seconds; the first measured CUDA run was slower than the next two. These results describe this small workload and runtime, not a universal speedup.
- CPU: AMD EPYC 7713; the server has two 64-core sockets, but this job used 4 PyTorch threads.
- GPU: NVIDIA RTX 6000 Ada Generation, 49140 MiB reported memory; driver 550.127.05.
- Runtime: Python 3.12, PyTorch 2.6.0+cu124, CUDA 12.4, Transformers 5.16.1, PEFT 0.20.0, Linux x86_64.
- Model: ESM3Di-35M, float32, evaluation mode, LoRA adapters merged into the backbone.
- Input: two ungapped amino-acid sequences, sorted by decreasing length, in one batch of 2.
- Exactly the same inputs and batch size were used for CPU and CUDA.
- One full-batch warmup, followed by three timed full-batch predictions per device.
- Prediction caches disabled; weights were already present on disk.
- Model loading excluded. Separately recorded loading time was 1.355539 seconds for CPU and 0.800689 seconds for CUDA; these are single loading observations, not cold-start benchmarks.
- Timed inference included tokenization, transfers, forward inference, finite-logit validation, argmax and return of prediction strings. CUDA was synchronized immediately before and after timing.
- Every timed output was compared with its warmup output; CUDA predictions were compared with CPU predictions.
- Output lengths were 960 and 418. The SHA256 of the two prediction strings joined by one LF
(PEPC first, PGK second, no final LF) was
0b7c1629398b71741744f837b67756734205a7029151a446751fa2c7a482ef02on both devices. - Slurm allocation: 4 CPUs, 10 GiB host memory, one GPU, 30-minute time limit. Other jobs were not modified. Queue time is excluded from the measurements.
| FASTA identifier | Protein | Residues |
|---|---|---|
PGK:Homo_sapiens_ENSG00000170950 |
Human phosphoglycerate kinase | 418 |
PEPC:Zea_mays_AB012228 |
Maize phosphoenolpyruvate carboxylase | 960 |
These are translations of the first records in the bundled
PGK CDS and
PEPC CDS,
with gaps removed and nonstandard residues represented by X, following the
prediction input handling. The terminal X in the supplied PGK input is retained.
The exact FASTA below uses one sequence line per record, LF line endings and a final LF.
Its SHA256 is 86bb22c676a04d3db6c749329bd6fcc7f37df883d96d925cc25e2a8ec47db921.
Exact test FASTA (1378 residues)
>PGK:Homo_sapiens_ENSG00000170950
MSLSKKLTLDKLDVRGKRVIMRVDFNVPMKKNQITNNQRIKASIPSIKYCLDNGAKAVVLMSHLGRPDGVPMPDKYSLAPVAVELKSLLGKDVLFLKDCVGAEVEKACANPAPGSVILLENLRFHVEEEGKGQDPSGKKIKAEPDKIEAFRASLSKLGDVYVNDAFGTAHRAHSSMVGVNLPHKASGFLMKKELDYFAKALENPVRPFLAILGGAKVADKIQLIKNMLDKVNEMIIGGGMAYTFLKVLNNMEIGASLFDEEGAKIVKDIMAKAQKNGVRITFPVDFVTGDKFDENAQVGKATVASGISPGWMGLDCGPESNKNHAQVVAQARLIVWNGPLGVFEWDAFAKGTKALMDEIVKATSKGCITVIGGGDTATCCAKWNTEDKVSHVSTGGGASLELLEGKILPGVEALSNMX
>PEPC:Zea_mays_AB012228
MPERHQSIDAQLRLLAPGKVSEDDKLVEYDALLVDRFLDILQDLHGPHLREFVQECYELSAEYENDRDEARLGELGSKLTSLPPGDSIVVASSFSHMLNLANLAEEVQIAHRRRIKLKRGDFADEASAPTESDIEETLKRLVSQLGKSREEVFDALKNQTVDLVFTAHPTQSVRRSLLQKHGRIRNCLRQLYAKDITADDKQELDEALQREIQAAFRTDEIRRTPPTPQDEMRAGMSYFHETIWKGVPKFLRRIDTALKNIGINERLPYNAPLIQFSSWMGGDRDGNPRVTPEVTRDVCLLARMMAANLYFSQIEDLMFELSMWRCSDELRIRADELHRSSRKAAKHYIEFWKQVPPNEPYRVILGDVRDKLYYTRERSRHLLTSGISEILEEATFTNVEQFLEPLELCYRSLCACGDKPIADGSLLDFLRQVSTFGLALVKLDIRQESDRHTDVLDSITTHLGIGSYAEWSEEKRQDWLLSELRGKRPLFGSDLPQTEETADVLGTFHVLAELPADCFGAYIISMATAPSDVLAVELLQRECHVKHPLRVVPLFEKLADLEAAPAAVARLFSIDWYMDRINGKQEVMIGYSDSGKDAGRLSAAWQMYKAQEELIKVAKHYGVKLTMFHGRGGTVGRGGGPTHLAILSQPPDTIHGSLRVTVQGEVIEHSFGEELLCFRTLQRYTAATLEHGMHPPISPKPEWRALMDEMAVVATKEYRSIVFQEPRFVEYFRSATPETEYGRMNIGSRPSKRKPSGGIESLRAIPWIFAWTQTRFHLPVWLGFGAAIKHIMQKDIRNIHILREMYNEWPFFRVTLDLLEMVFAKGDPGIAAVYDKLLVADDLQSFGEQLRKNYEETKELLLQVAGHKDVLEGDPYLKQRLRLRESYITTLNVCQAYTLKRIRDPSFQVSPQPPLSKEFTDESQPAELVQLNQQSEYAPGLEDTLILTMKGIAAGMQNTG
- Trained checkpoint:
cactuskid13/esm2small_3di, revision2227cb08ffa533bc1a7f2b968fd6316f287134d1, fileepoch_3.pt. - Checkpoint SHA256:
044489c3bf8ffa1067dd2e71edf1750b48a2d52d73018ae4908d98ddebbb0587. - Backbone configuration/tokenizer:
facebook/esm2_t12_35M_UR50D, revision6fbf070e65b0b7291e7bbcd451118c216cff79d8. The trained checkpoint already contains the backbone weights. - Tested implementation file SHA256 values (remote copy verified against the local worktree):
| File | SHA256 |
|---|---|
csubst/structural_prediction.py |
3abb859751cbbe80dfd689e90461c7a8a3035cbca97a591b7c57865d1a076948 |
csubst/structural_alphabet.py |
3d59a124836076be2335cee5c9bdb1c9f1a762fddb63892d42e7fca9e2da7752 |
csubst/sequence.py |
f727a1bbaa72af814aef31ebfc9c0057b2b8295b73f2e51d49f7bfb8b5e52829 |
csubst/recoding_config.py |
30a174ca9a5f3d55245584676cad98fffb1beef81f367331f5be9005e4ab2856 |
This is a model-inference timing comparison, not whole-command or end-to-end ancestral reconstruction timing. It does not measure structure-reference prediction accuracy or justify biological equivalence to ProstT5. The released ESM3Di checkpoint was trained on viral BFVD data; these PGK/PEPC inputs test execution and timing on two other proteins. Larger datasets, sequence lengths, batch sizes, CPU thread counts and GPU/runtime combinations can behave differently.
The earlier functional-audit timings are deliberately excluded: its CUDA-only extra long-sequence checks made the ESM3Di CPU and GPU timed workloads unequal.