GlubLM v0.3.0
Architecture upgrade: 18M -> 35M parameters, 48 -> 96 token context.
Architecture
- 36.1M params (d_model=640, n_heads=10, ffn_hidden=1280)
- 96-token context window (was 48)
- Same RoPE + SwiGLU + RMSNorm stack
- ONNX uint8: ~40MB (browser-friendly)
Results
- Final loss: 1.19 (was 1.70)
- Significantly cleaner grammar, better comprehension
- Same goldfish personality (persona is in the data, not the params)
What improved
- Less grammatical garbling (~5% vs ~15%)
- Better understanding of user messages
- More coherent responses
- Longer generation window without truncation