Replies: 3 comments 2 replies
|
I can attempt that tomorrow. |
|
Run 6 baselines on 8x RTX PRO 6000s
|
|
Worth weighing those two numbers against the spread the leaderboard already documents. For Run 4, karpathy ran the identical command 7 times on one commit and got CORE 0.25373, 0.2584, 0.25489, 0.2568, 0.25732, 0.26765, 0.25119: max-min 0.01646, sample stdev 0.0052 (dev/LEADERBOARD.md). The gap here, 0.2603 - 0.2516 = 0.0087, is about 1.2 stdev of the difference between two single runs. And Run 6's published 0.262634 is itself an average of 5 runs, so 0.2603 at val bpb is what the same doc points at for exactly this ("has less noise than CORE"), and there master is 0.718948 against 0.719839, i.e. 0.0009 lower. No regression shows up on the smooth metric. So one more 8xH100 run probably won't settle it either. Telling a real regression apart from shuffle/nondeterminism noise at this size takes several runs per commit. |
Uh oh!
There was an error while loading. Please reload this page.
I got a DM on LinkedIn from someone who wasn't able to reliably reproduce the run 6 baseline CORE with
master.Has anyone run one recently? If nobody has, I'll try and get a hold of an 8xH100 to verify sometime in the next few weeks.
All reactions