Skip to content

ESM3Di CPU GPU benchmark

Kenji Fukushima edited this page Sep 9, 2026 · 2 revisions

ESM3Di CPU/GPU benchmark

Report version 1.0.1 (release-status clarification; measurements unchanged). Measured on audrey1 on 2026-09-10 (JST), Slurm job 25616.

For the two proteins below (1378 residues total), median ESM3Di-35M inference time was 0.901931 seconds on CPU and 0.015422 seconds on CUDA: a 58.48× CPU/GPU elapsed-time ratio. Predictions were identical across CPU and CUDA and across all measured repetitions.

At measurement time, this used an unreleased CSUBST development worktree, based on commit 286ac2b, with local 3Di predictor changes. It is not a benchmark of the published base commit alone. The tested predictor source hashes below also match CSUBST 1.15.0, which has since been released on GitHub. Distribution availability is tracked in the installation guide.

Timing results

All times are seconds for the entire batch of two sequences.

Device Run 1 Run 2 Run 3 Median
CPU (4 threads) 0.901879553 0.905969813 0.901930879 0.901930879
CUDA (RTX 6000 Ada) 0.022375485 0.015422144 0.015390404 0.015422144

The ratio is CPU median / CUDA median, not a ratio of model-loading times. Only three repetitions were measured. CUDA values ranged from 0.015390 to 0.022375 seconds; the first measured CUDA run was slower than the next two. These results describe this small workload and runtime, not a universal speedup.

Hardware and method

  • CPU: AMD EPYC 7713; the server has two 64-core sockets, but this job used 4 PyTorch threads.
  • GPU: NVIDIA RTX 6000 Ada Generation, 49140 MiB reported memory; driver 550.127.05.
  • Runtime: Python 3.12, PyTorch 2.6.0+cu124, CUDA 12.4, Transformers 5.16.1, PEFT 0.20.0, Linux x86_64.
  • Model: ESM3Di-35M, float32, evaluation mode, LoRA adapters merged into the backbone.
  • Input: two ungapped amino-acid sequences, sorted by decreasing length, in one batch of 2.
  • Exactly the same inputs and batch size were used for CPU and CUDA.
  • One full-batch warmup, followed by three timed full-batch predictions per device.
  • Prediction caches disabled; weights were already present on disk.
  • Model loading excluded. Separately recorded loading time was 1.355539 seconds for CPU and 0.800689 seconds for CUDA; these are single loading observations, not cold-start benchmarks.
  • Timed inference included tokenization, transfers, forward inference, finite-logit validation, argmax and return of prediction strings. CUDA was synchronized immediately before and after timing.
  • Every timed output was compared with its warmup output; CUDA predictions were compared with CPU predictions.
  • Output lengths were 960 and 418. The SHA256 of the two prediction strings joined by one LF (PEPC first, PGK second, no final LF) was 0b7c1629398b71741744f837b67756734205a7029151a446751fa2c7a482ef02 on both devices.
  • Slurm allocation: 4 CPUs, 10 GiB host memory, one GPU, 30-minute time limit. Other jobs were not modified. Queue time is excluded from the measurements.

Test data

FASTA identifier Protein Residues
PGK:Homo_sapiens_ENSG00000170950 Human phosphoglycerate kinase 418
PEPC:Zea_mays_AB012228 Maize phosphoenolpyruvate carboxylase 960

These are translations of the first records in the bundled PGK CDS and PEPC CDS, with gaps removed and nonstandard residues represented by X, following the prediction input handling. The terminal X in the supplied PGK input is retained. The exact FASTA below uses one sequence line per record, LF line endings and a final LF. Its SHA256 is 86bb22c676a04d3db6c749329bd6fcc7f37df883d96d925cc25e2a8ec47db921.

Exact test FASTA (1378 residues)
>PGK:Homo_sapiens_ENSG00000170950
MSLSKKLTLDKLDVRGKRVIMRVDFNVPMKKNQITNNQRIKASIPSIKYCLDNGAKAVVLMSHLGRPDGVPMPDKYSLAPVAVELKSLLGKDVLFLKDCVGAEVEKACANPAPGSVILLENLRFHVEEEGKGQDPSGKKIKAEPDKIEAFRASLSKLGDVYVNDAFGTAHRAHSSMVGVNLPHKASGFLMKKELDYFAKALENPVRPFLAILGGAKVADKIQLIKNMLDKVNEMIIGGGMAYTFLKVLNNMEIGASLFDEEGAKIVKDIMAKAQKNGVRITFPVDFVTGDKFDENAQVGKATVASGISPGWMGLDCGPESNKNHAQVVAQARLIVWNGPLGVFEWDAFAKGTKALMDEIVKATSKGCITVIGGGDTATCCAKWNTEDKVSHVSTGGGASLELLEGKILPGVEALSNMX
>PEPC:Zea_mays_AB012228
MPERHQSIDAQLRLLAPGKVSEDDKLVEYDALLVDRFLDILQDLHGPHLREFVQECYELSAEYENDRDEARLGELGSKLTSLPPGDSIVVASSFSHMLNLANLAEEVQIAHRRRIKLKRGDFADEASAPTESDIEETLKRLVSQLGKSREEVFDALKNQTVDLVFTAHPTQSVRRSLLQKHGRIRNCLRQLYAKDITADDKQELDEALQREIQAAFRTDEIRRTPPTPQDEMRAGMSYFHETIWKGVPKFLRRIDTALKNIGINERLPYNAPLIQFSSWMGGDRDGNPRVTPEVTRDVCLLARMMAANLYFSQIEDLMFELSMWRCSDELRIRADELHRSSRKAAKHYIEFWKQVPPNEPYRVILGDVRDKLYYTRERSRHLLTSGISEILEEATFTNVEQFLEPLELCYRSLCACGDKPIADGSLLDFLRQVSTFGLALVKLDIRQESDRHTDVLDSITTHLGIGSYAEWSEEKRQDWLLSELRGKRPLFGSDLPQTEETADVLGTFHVLAELPADCFGAYIISMATAPSDVLAVELLQRECHVKHPLRVVPLFEKLADLEAAPAAVARLFSIDWYMDRINGKQEVMIGYSDSGKDAGRLSAAWQMYKAQEELIKVAKHYGVKLTMFHGRGGTVGRGGGPTHLAILSQPPDTIHGSLRVTVQGEVIEHSFGEELLCFRTLQRYTAATLEHGMHPPISPKPEWRALMDEMAVVATKEYRSIVFQEPRFVEYFRSATPETEYGRMNIGSRPSKRKPSGGIESLRAIPWIFAWTQTRFHLPVWLGFGAAIKHIMQKDIRNIHILREMYNEWPFFRVTLDLLEMVFAKGDPGIAAVYDKLLVADDLQSFGEQLRKNYEETKELLLQVAGHKDVLEGDPYLKQRLRLRESYITTLNVCQAYTLKRIRDPSFQVSPQPPLSKEFTDESQPAELVQLNQQSEYAPGLEDTLILTMKGIAAGMQNTG

Model and source provenance

  • Trained checkpoint: cactuskid13/esm2small_3di, revision 2227cb08ffa533bc1a7f2b968fd6316f287134d1, file epoch_3.pt.
  • Checkpoint SHA256: 044489c3bf8ffa1067dd2e71edf1750b48a2d52d73018ae4908d98ddebbb0587.
  • Backbone configuration/tokenizer: facebook/esm2_t12_35M_UR50D, revision 6fbf070e65b0b7291e7bbcd451118c216cff79d8. The trained checkpoint already contains the backbone weights.
  • Tested implementation file SHA256 values (remote copy verified against the local worktree):
File SHA256
csubst/structural_prediction.py 3abb859751cbbe80dfd689e90461c7a8a3035cbca97a591b7c57865d1a076948
csubst/structural_alphabet.py 3d59a124836076be2335cee5c9bdb1c9f1a762fddb63892d42e7fca9e2da7752
csubst/sequence.py f727a1bbaa72af814aef31ebfc9c0057b2b8295b73f2e51d49f7bfb8b5e52829
csubst/recoding_config.py 30a174ca9a5f3d55245584676cad98fffb1beef81f367331f5be9005e4ab2856

Scope

This is a model-inference timing comparison, not whole-command or end-to-end ancestral reconstruction timing. It does not measure structure-reference prediction accuracy or justify biological equivalence to ProstT5. The released ESM3Di checkpoint was trained on viral BFVD data; these PGK/PEPC inputs test execution and timing on two other proteins. Larger datasets, sequence lengths, batch sizes, CPU thread counts and GPU/runtime combinations can behave differently.

The earlier functional-audit timings are deliberately excluded: its CUDA-only extra long-sequence checks made the ESM3Di CPU and GPU timed workloads unequal.

Clone this wiki locally