Repository navigation
cuequivariance vs stock e3nn: a flat sub-unity region below break-even, and where break-even lands (MACE-MP-0, A5000) #180
liulangdog1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
While profiling small-system MLIP workloads I ended up characterising where
cuequivariancestarts to pay for itself against stocke3nn. The shape turned out to be more consistent than the crossing point, so I'm posting both. Everything below is a single fused forward with forces on a pre-built batch — kernel-level, not an end-to-end workflow number.The shape: a flat penalty, then a rise
Below break-even the ratio is not gradually degrading, it is constant. Small/float32 sits at 0.72–0.73 across every batch size from 1 through 48, then lifts. The same flat-then-rise appears in all six model × precision combinations tested.
At batch 1,
cuequivarianceis 0.62–0.73× of stocke3nnin all six — a narrow band across a 4× parameter range and both precisions.That is consistent with a fixed per-launch kernel overhead: constant in relative terms until there is enough work per launch to hide it.
Ratio (cuequivariance ÷ stock e3nn), float32:
float64:
Break-even, as brackets
Reported as the pair of measured batch sizes that straddle 1.0. Nothing interpolated — the sweep spacing does not support a single value.
Monotonic in model size in both precisions.
What moves it, and by how much
The crossing is not portable. Three things move it:
Note that parameter count is a poor abscissa here. All three checkpoints report
num_interactions: 2, so they differ in angular momentum rather than depth: small → medium is +22% parameters but moves the crossing by roughly 1.5 brackets, while medium → large moves it much further. I did not capturehidden_irrepson this run, so I can't yet plot against a proper work-per-structure axis.One thing I could not reconcile
A separate run on 32-atom fcc Cu, using a different harness, put the small model's crossing near batch 17 rather than b96–128. Recasting the crossing as total atoms per forward pass rather than batch count narrows the discrepancy from roughly 7× to under 4×, which suggests atoms-per-batch is the better abscissa — but the two runs still do not fully reconcile, and I have not established why. Flagging it rather than smoothing it over.
Method
RTX A5000 (sm86, 23 GB), driver 575.57.08, CUDA 12.6. torch 2.12.0+cu126, mace-torch 0.3.16, cuequivariance 0.10.0. Conversion via
mace.cli.convert_e3nn_cueq.runwith"convert": truerecorded. Ethanol (C₂H₆O, 9 atoms) from ASE's G2 database. MACE-MP-0 small/medium/large. Best of 5 after 3 warm-ups. Single node, single card, single driver throughout.Instrument check. Before any timing was read,
torch.opswas confirmed to expose['cuequivariance', 'cuequivariance_ops']withfused_tensor_productpresent. Worth stating because the namespace moved: on cuequivariance 0.6.0 the ops live undertorch.ops.cuequivariance_opsand the forward isfused_tensor_product_fwd, so the obvious check false-negatives on older versions and you can silently benchmark a fallback path.Cross-script reproduction. The float64 brackets above (small b8→b10, medium b3→b4) match an independent run with a different harness on the same system (small b8→b16, medium b2→b4).
One question
The contributing guide says direct code contributions are paused during the public beta. Is that still current? If a reproducible benchmark script for this would be useful to have in the repo I'm happy to open one; otherwise I'll leave the numbers here.
All reactions