Skip to content

Jeff v1.2: nine adapters, and Jeff in front of your 27B

Latest

Choose a tag to compare

@firelex firelex released this 01 Oct 12:03
· 1 commit to main since this release

Jeff v1.2: a cleaner base, nine adapters, clients

Fast, calibrated decisions for local agents: one base, swappable adapters.

Put Jeff in front of your 27B

Jeff + its adapters answer first; only when Jeff is unsure does the query go on to Qwen3.8-27B. Across 8 adapters*,
same test rows both ways, on an Apple M4 Max:

Qwen3.8-27B alone Jeff + adapters, the 27B only when unsure
Accuracy (mean of 8 adapters*) 86.6% 95.3%
Time per decision (mean) 8.1 s 0.25 s: 38× faster
Wrong answers 13.4% 4.7%
Memory 28.6 GB +1.96 GB for Jeff with all nine adapters loaded

On the five decisions an inbox agent makes for every message alone: 87.7% → 95.7%, 39× faster. The 27B ran in 8-bit
with step-by-step reasoning off; each adapter's threshold was chosen on separate calibration rows. Full method:
README and the
reference app.

*Emotion is left out of the averages (27 emotions or neutral in short comments, where even human labels disagree):
Jeff + adapter 60.6% against the 27B's 35.6%. Including it, the average across all nine is 91.4% against
80.9%.

Jeff-Qwen3.5-0.8B v1.2: more honest, not smarter

We found that v1.1's training data overlapped some of our own test sets, and that parts of it let a model score by
answer order or answer length instead of by content. v1.2 is the same recipe on cleaned data (284,747 questions):

  • Leaks removed. 5,279 training questions overlapped our long-document test (all 4,309 MAUD questions among them:
    MAUD splits by question, not by contract), 4,358 our first long-list test, 203 the voice test and 66 the benchmark
    panel. All are gone, and the leak check now runs before every training run.
  • Shortcuts removed. Answer length, answer letter and option count no longer give answers away in the synthetic
    data; synthetic yes/no families are 50/50.
  • Voice navigation moved out of the base and into its own adapter.
Test v1.1 v1.2
Benchmark panel, 5 benchmarks (4,599 questions) 79.1% 78.7%
Calibration error on the panel (ECE) 0.021 0.028
Long lists v2, 20 to 254 options (1,886; new, clean) not measured 93.1%
Long documents (2,009) 82.8% 66.1%: v1.1 was inflated by leaked MAUD documents
Voice navigation (3,324) 95.0% 90.2%: now zero-shot
JevBench hard tier (105) 46.7% 44.8%

The benchmark score is level and calibration slightly worse; the long-document drop is the leak coming out. v1.1 stays
on Hugging Face as revision v1.1. Jeff-Qwen3.5-2B is retrained on the same data too: panel 82.0% → 81.7%, calibration
error 0.026 → 0.021, long documents 85.8% → 65.6% (the same leak), JevBench hard 57.1% (unchanged), long lists v2 93.6%.

Tried and dropped: a 0.8B distilled from the 2B scored the same as v1.2 on every test, so it is not released.

Nine LoRA adapters

About 41 MB each, on the v1.2 base. Pick the ones you need: one server loads the base once and the adapters beside it,
and each request picks one by name in "model" (or none, for plain Jeff). Accuracy on each adapter's full held-out
test set, included in its repository with its calibration set:

Adapter Test set Qwen3.5-0.8B untrained Jeff v1.2 alone Jeff v1.2 + adapter
legal-clauses LEDGAR test split, 100 clause types (9,895 rows) 12.5% 66.0% 85.7%
support-intents Bitext, HWU64 and SNIPS (5,577 rows) 33.9% 85.1% 96.8%
emotion GoEmotions, 27 emotions plus neutral (5,408 rows) 12.5% 32.2% 60.6%
spam SMS and email (3,603 rows) 59.4% 72.3% 98.4%
ground held-out articles and documents (4,160 rows) 28.7% 49.0% 97.0%
guard held-out applications (6,552 rows) 43.7% 46.9% 98.4%
tools held-out agents (5,157 rows) 18.0% 57.8% 97.9%
nav held-out apps (3,300 rows) 12.6% 23.8% 97.0%
triage held-out organisations (7,256 rows) 44.4% 67.1% 91.8%

Part of the data of ground, guard, tools, nav and triage was written by Qwen3.8-Max, a hosted, closed-weight
model, through Alibaba Cloud's DashScope API, and part was checked by hosted Qwen3.8-Flash; each card gives the exact
counts. legal-clauses, support-intents, emotion and spam use public data only.

Code

  • jeff-serve with JEFF_ADAPTERS: several adapters on one base (PyTorch and MLX), chosen per request, added or
    replaced without a restart (POST /v1/adapters/reload).
  • jeff-train --lora-rank ... trains an adapter; jeff-evaluate and jeff-latency accept adapter folders.
  • Python client (from jeff import Client, standard library only) and TypeScript client (clients/typescript).
  • Answer twice: "orders": 2 answers each question again with the options reversed and averages the two.
  • Training accepts score questions (ordered levels).

Try it

git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --revision v1.2 --local-dir Jeff-Qwen3.5-0.8B-v1.2
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B-guard --local-dir adapters/guard
JEFF_CHECKPOINT=Jeff-Qwen3.5-0.8B-v1.2 JEFF_ADAPTERS=adapters PORT=8765 uv run --no-default-groups --extra lora jeff-serve