README refresh: the evals section now reports real model measurements — open-weights clef-flash 9B (float16, Kaggle 2x T4) on the committed gold dataset: chunk accuracy 0.712, kept F1 0.758, 38.4% of context tokens removed, scoring latency 1,274 ms p50 self-hosted. Reproducible from the public clef-compactor-evals kernel.
Also ships evals/backends.py (local-weights backend) and --mode local in the eval harness. No API changes.