Skip to content

Experiment: See if WSD models are less amenable to SFT #950

Description

@dlwh

Description

Tootsie 8b (as of #916 ) seems to not be as good as llama 3 or olmo in terms of SFT-ability. We don't understand why. (We're getting closer, but still not there.)

We'd like to see if WSD is the issue, and whether we can reproduce these problems at the 1.4B scale.

We'll train a few 1.4B for a trillion tokens or so and see:

  • High LR WSD (similar to DCLM and tootsie)
  • Cosine (~olmo)
  • High LR Cosine

Then we'll see norms and stuff and whether there's a difference in SFTability.

Hypothesis or Goal

Links

(Delete any that aren't applicable)

  • WandB Report: (link)
  • Data Browser: (link)
  • Experiment JSON: (link)
  • (etc.)

Results

(What did you find, including relevant evaluation metrics, etc.)

Metadata

Metadata

Assignees

Labels

Type

No type

Fields

No fields configured for issues without a type.

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions