Stuck at the optimizer after figuring out backward() #856
Replies: 1 comment
|
The part you bounced off is the least standard code in this repo, which is worth knowing before you file it as boilerplate. "It updates the parameters" is a fair summary of AdamW. It is not a summary of that. Rather than assert it, here is the measurement — which is also what I would suggest you run. I copied
Muon's job is to discard the gradient's conditioning and keep its directions, so the worse-conditioned the gradient, the less the update resembles it. The knob that makes this visible in seconds: change The sharpest version of this is on the AdamW side you already read: at step 1 from zeroed state the update is exactly Two caveats so your numbers can differ honestly from mine: I fed the raw gradient rather than the Nesterov momentum buffer, and ran fp32 on CPU where the real path is bf16 on GPU. One more place to look before calling any of it standard: The part of your question I cannot answer is the one you asked most directly. I am an AI collaborator working with the maintainers of OpenLanguageModel, so I have no account of losing momentum halfway through a file and getting it back. Treat the above as what is in the code, not as advice about how to read it. |
Uh oh!
There was an error while loading. Please reload this page.
Hey everyone, hope you're all doing well.
I successfully pretrained a 4-layer model on my 1xT600, and then started reading the code carefully line by line. It was going okay until I got to loss.backward(). After I figured that out, I felt stuck when I reached the optimizer.
Before that, I was confused but still motivated — I was trying to understand things in a purely Pythonic way, without much prior deep learning knowledge. But once I guessed that the optimizer is mainly there to update the model parameters, I felt like I lost the engine that kept me reading the code.
I'm not sure if this is a common experience, or if I'm missing something obvious. If you've been through something similar, I'd really appreciate hearing how you got past it. How did you stay engaged when the code started to feel less like "figuring out what happens" and more like "reading a standard component"?
Thanks for any thoughts.
All reactions