TUFF 6.1.0
TUFF 6.1.0 extends the efficient GPU sampler to Flash Next top-k 20 and MiniMax top-k 40, and adds repeatable inference comparisons and clearer expert-cache diagnostics.
- Separate demand and prediction histories. Report first demand hits on prefetched records, unused evictions and exposed waits while retaining the validated eviction ranking.
- Retain the synchronized expert streamer for background lookahead, with a startup-only off switch for comparisons.
- Account for the extra cache metadata in shared app, CLI and server memory budgets. Keep qualified cache, grouping and chunk defaults.
- Record short/long, greedy/default sampled and cold/warm comparisons, resolved settings, model identities and available machine-state observations.
Isolated sampler wall medians improved from 23.20 to 1.27 ms for Flash Next and 58.87 to 1.03 ms for MiniMax. Full-inference results were mixed. Demand-only eviction, unused-prefetch priority and grouping experiments were removed; aging was not qualified end to end. No general model speedup is claimed.
Validation: 1,682 canonical tests across 12 targets, 29 harness/reporting regressions, repository checks and release build/archive checks. All nine text models and six image companions passed packaged CLI checks, with nine app-service and 18 HTTP-server passes. The warm comparison passed 96 requests with matching visible output in all 48 pairs. Four sleep-interrupted timing observations remain labeled, with awake replacements. All model work ran serially on one 16 GB M2 MacBook Air; other hardware was not tested.
The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed. See the release validation and model report for settings, all repetitions and limitations.
AI assistance: OpenAI Codex helped implement, review, test and document these changes. Results came from the checks described above.