Description
Tootsie 8b (as of #916 ) seems to not be as good as llama 3 or olmo in terms of SFT-ability. We don't understand why. (We're getting closer, but still not there.)
We'd like to see if WSD is the issue, and whether we can reproduce these problems at the 1.4B scale.
We'll train a few 1.4B for a trillion tokens or so and see:
- High LR WSD (similar to DCLM and tootsie)
- Cosine (~olmo)
- High LR Cosine
Then we'll see norms and stuff and whether there's a difference in SFTability.
Hypothesis or Goal
Links
(Delete any that aren't applicable)
- WandB Report: (link)
- Data Browser: (link)
- Experiment JSON: (link)
- (etc.)
Results
(What did you find, including relevant evaluation metrics, etc.)
Description
Tootsie 8b (as of #916 ) seems to not be as good as llama 3 or olmo in terms of SFT-ability. We don't understand why. (We're getting closer, but still not there.)
We'd like to see if WSD is the issue, and whether we can reproduce these problems at the 1.4B scale.
We'll train a few 1.4B for a trillion tokens or so and see:
Then we'll see norms and stuff and whether there's a difference in SFTability.
Hypothesis or Goal
Links
(Delete any that aren't applicable)
Results
(What did you find, including relevant evaluation metrics, etc.)