autoresearch-xv — one vendor-agnostic source (NVIDIA+AMD) for measuring the cross-vendor efficiency gap #641
crivayne
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Most forks specialize to one platform. This one goes the other way: it keeps
train.pyas a single vendor-agnostic source that runs unchanged on NVIDIA and AMD, and uses that source to study how much portability costs — or does not cost — relative to per-vendor tuning. It includes explicit RDNA4/ROCm support for the RX 9070 XT (gfx1201) on native Windows and Linux, while currently recommending Linux for quality-sensitive training.How it works
try/excepton non-ROCm systems; ROCm and FA3 load failures use full-causal PyTorch SDPA. MFU uses a device-name peak-FLOPS table. Escape hatches includeAR_FORCE_SDPA=1for controlled SDPA comparisons on NVIDIA,AR_NO_COMPILE=1, and several additionalAR_*diagnostic knobs.torch.compile, TunableOp, andDEVICE_BATCH_SIZE=32— reachesval_bpb1.5039, ~190K tok/s, and 47.3% repo-computed MFU on the RX 9070 XT. This slightly edges the published MI308X eager reference (val_bpb1.5208, ~161K tok/s) within the same fixed 5-minute training budget, though the hardware and software stacks differ.val_bpb1.1177 through budget-specialized optimization. Those results are kept on a separate axis from the fixed-architecture standard-recipe comparison.A finding worth flagging for RDNA4 users
In my RX 9070 XT tests, Windows-native ROCm 7.2.1 showed systematically worse bf16 training convergence than the Linux ROCm stack on the same GPU. The OS-stack gap was roughly 10× the observed seed-to-seed standard deviation across three Windows-side seeds. Windows-side diagnostic changes — including fp32 SDPA, fp32 Muon orthogonalization, and BLAS-backend selection — did not remove it.
For this configuration, I exclude Windows-native results from training-quality comparisons and use Linux ROCm for quality-sensitive training. Windows remains useful for inference and tooling. An upstream report is in preparation; confirmation on the latest TheRock nightly is still pending.
How this differs from the existing
autoresearch-rocmentryThe scopes are complementary.
JKeller45/autoresearch-rocmtargets AMD and officially supports/tests the ROCm-on-WSL2 path; it hard-fails without HIP and intentionally uses eager SDPA withtorch.compileand TF32 disabled. This fork instead keeps one NVIDIA+AMDtrain.py, runstorch.compileplus TunableOp on native Linux, and focuses on measuring cross-vendor behavior from a shared source.Repo: https://github.com/crivayne/autoresearch-xv
All reactions