Skip to content

Releases: firelex/jeff

v1.1: up to 254 options

Choose a tag to compare

@firelex firelex released this 29 Sep 18:28

Jeff-Qwen3.5-0.8B and Jeff-Qwen3.5-2B v1.1 accept up to 254 options per question (v1.0: 26).

  • Long lists: 32,000 new training questions with 20 to 254 options (code-built lists, and the MASSIVE and CLINC150 training splits). Long-list test: 0.8B 40.3% → 94.7%; 2B 95.2%. Thanks to @puhuk for the report (#1).
  • Final checkpoint: v1.1 publishes the end-of-epoch checkpoint; development-loss selection had picked an early, less settled checkpoint.
  • Calibration: error 0.049 → 0.021 (0.8B), 0.028 → 0.026 (2B). JevBench hard tier: 46.7% (0.8B), 57.1% (2B, up from 53.3%).
  • Benchmarks: 0.8B unchanged at 79.1%; 2B 83.1% → 82.0%, mostly JudgeBench (64.6% → 59.4%), which sits near chance at this size.
  • Serving-only install: uv sync --no-default-groups (#2, thanks to @WavesMan).
  • New base option: Phi-4-mini-instruct can now be fine-tuned (#4, thanks to @mikeatlas).
  • MASSIVE and CLINC150 results are no longer zero-shot. Jeff-Gemma4-E2B stays at v1.0.

Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B. v1.0 stays available on Hugging Face as revision v1.0.