Jeff v1.2: a cleaner base, nine adapters, clients
Fast, calibrated decisions for local agents: one base, swappable adapters.
Put Jeff in front of your 27B
Jeff + its adapters answer first; only when Jeff is unsure does the query go on to Qwen3.8-27B. Across 8 adapters*,
same test rows both ways, on an Apple M4 Max:
| Qwen3.8-27B alone | Jeff + adapters, the 27B only when unsure | |
|---|---|---|
| Accuracy (mean of 8 adapters*) | 86.6% | 95.3% |
| Time per decision (mean) | 8.1 s | 0.25 s: 38× faster |
| Wrong answers | 13.4% | 4.7% |
| Memory | 28.6 GB | +1.96 GB for Jeff with all nine adapters loaded |
On the five decisions an inbox agent makes for every message alone: 87.7% → 95.7%, 39× faster. The 27B ran in 8-bit
with step-by-step reasoning off; each adapter's threshold was chosen on separate calibration rows. Full method:
README and the
reference app.
*Emotion is left out of the averages (27 emotions or neutral in short comments, where even human labels disagree):
Jeff + adapter 60.6% against the 27B's 35.6%. Including it, the average across all nine is 91.4% against
80.9%.
Jeff-Qwen3.5-0.8B v1.2: more honest, not smarter
We found that v1.1's training data overlapped some of our own test sets, and that parts of it let a model score by
answer order or answer length instead of by content. v1.2 is the same recipe on cleaned data (284,747 questions):
- Leaks removed. 5,279 training questions overlapped our long-document test (all 4,309 MAUD questions among them:
MAUD splits by question, not by contract), 4,358 our first long-list test, 203 the voice test and 66 the benchmark
panel. All are gone, and the leak check now runs before every training run. - Shortcuts removed. Answer length, answer letter and option count no longer give answers away in the synthetic
data; synthetic yes/no families are 50/50. - Voice navigation moved out of the base and into its own adapter.
| Test | v1.1 | v1.2 |
|---|---|---|
| Benchmark panel, 5 benchmarks (4,599 questions) | 79.1% | 78.7% |
| Calibration error on the panel (ECE) | 0.021 | 0.028 |
| Long lists v2, 20 to 254 options (1,886; new, clean) | not measured | 93.1% |
| Long documents (2,009) | 82.8% | 66.1%: v1.1 was inflated by leaked MAUD documents |
| Voice navigation (3,324) | 95.0% | 90.2%: now zero-shot |
| JevBench hard tier (105) | 46.7% | 44.8% |
The benchmark score is level and calibration slightly worse; the long-document drop is the leak coming out. v1.1 stays
on Hugging Face as revision v1.1. Jeff-Qwen3.5-2B is retrained on the same data too: panel 82.0% → 81.7%, calibration
error 0.026 → 0.021, long documents 85.8% → 65.6% (the same leak), JevBench hard 57.1% (unchanged), long lists v2 93.6%.
Tried and dropped: a 0.8B distilled from the 2B scored the same as v1.2 on every test, so it is not released.
Nine LoRA adapters
About 41 MB each, on the v1.2 base. Pick the ones you need: one server loads the base once and the adapters beside it,
and each request picks one by name in "model" (or none, for plain Jeff). Accuracy on each adapter's full held-out
test set, included in its repository with its calibration set:
| Adapter | Test set | Qwen3.5-0.8B untrained | Jeff v1.2 alone | Jeff v1.2 + adapter |
|---|---|---|---|---|
| legal-clauses | LEDGAR test split, 100 clause types (9,895 rows) | 12.5% | 66.0% | 85.7% |
| support-intents | Bitext, HWU64 and SNIPS (5,577 rows) | 33.9% | 85.1% | 96.8% |
| emotion | GoEmotions, 27 emotions plus neutral (5,408 rows) | 12.5% | 32.2% | 60.6% |
| spam | SMS and email (3,603 rows) | 59.4% | 72.3% | 98.4% |
| ground | held-out articles and documents (4,160 rows) | 28.7% | 49.0% | 97.0% |
| guard | held-out applications (6,552 rows) | 43.7% | 46.9% | 98.4% |
| tools | held-out agents (5,157 rows) | 18.0% | 57.8% | 97.9% |
| nav | held-out apps (3,300 rows) | 12.6% | 23.8% | 97.0% |
| triage | held-out organisations (7,256 rows) | 44.4% | 67.1% | 91.8% |
Part of the data of ground, guard, tools, nav and triage was written by Qwen3.8-Max, a hosted, closed-weight
model, through Alibaba Cloud's DashScope API, and part was checked by hosted Qwen3.8-Flash; each card gives the exact
counts. legal-clauses, support-intents, emotion and spam use public data only.
Code
jeff-servewithJEFF_ADAPTERS: several adapters on one base (PyTorch and MLX), chosen per request, added or
replaced without a restart (POST /v1/adapters/reload).jeff-train --lora-rank ...trains an adapter;jeff-evaluateandjeff-latencyaccept adapter folders.- Python client (
from jeff import Client, standard library only) and TypeScript client (clients/typescript). - Answer twice:
"orders": 2answers each question again with the options reversed and averages the two. - Training accepts
scorequestions (ordered levels).
Try it
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --revision v1.2 --local-dir Jeff-Qwen3.5-0.8B-v1.2
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B-guard --local-dir adapters/guard
JEFF_CHECKPOINT=Jeff-Qwen3.5-0.8B-v1.2 JEFF_ADAPTERS=adapters PORT=8765 uv run --no-default-groups --extra lora jeff-serve