Repository navigation
Trained weights for the two models described in the README and paper/main.pdf (character-level enwik8, vocabulary 481).
| file | params | width | bpc at ctx 128 / 512 | sha256 |
|---|---|---|---|---|
| saryu_25m.pt | 25,183,861 | 744 | 1.535 / 1.433 | 2a7e420605540454d637d7af77bc71d4bce883ebcdbf89a8ca3929115439c6e4 |
| saryu_v4b_last.pt | 5,016,097 | 320 (value embedding) | 1.671 / 1.586 | 91815b36703ba4ee1ac4418b182b3d820c2593d277cf89253ca10f477bf828b4 |
Scores are bits per character on the last 128 characters of 64 evaluation windows, with identical
text scored at every context length. Both models stop using context at about 512 characters, so
1.433 and 1.586 are also their scores at 2048 and 8192 (evidence/results/effective_context.txt).
These numbers replace the ones published with this release on 2026-09-15, which were measured with
window positions drawn separately per length and are not comparable.
Download into checkpoints/: gh release download v0.1 --repo varun29ankuS/Saryu-RNN -D checkpoints
Loading uses torch.load(weights_only=False); only load checkpoints from sources you trust. Apache-2.0, as the code.