Skip to content

TUFF v6.0.2

Choose a tag to compare

@rexmhall09 rexmhall09 released this 01 Oct 05:25

TUFF 6.0.2 corrects inference memory planning and performance reporting.

  • Count MoE expert-cache slots across every layer, including GPT-OSS. Cross-check all six MoE layouts against their manifests and actual allocations.
  • Include chunk-dependent prefill scratch, sliding-window rings and conservative growth reserves in the shared memory plan used by the app, CLI and server.
  • Report demand and prefetched expert records, logical bytes and failures separately. Measure exposed prefetch waits and remove the unsupported GPU-wait residual label.
  • Keep every benchmark repetition with median and spread, resolved settings and available machine-state observations. Restore all measured README benchmarks and update the emoji comparison with sourced capabilities and current TUFF features.

Validation: 1,662 canonical tests across 12 targets, 11 benchmark-reporting regression tests, repository checks, release build and archive checks. The packaged CLI passed 27 text attempts across all nine supported models and all six image companions. Flash Next and Gemma 26B also passed packaged app decode-service and HTTP-server checks. All real-model work ran serially on a 16 GB M2 MacBook Air.

These are correctness smoke checks, not general model-quality or sustained-performance qualification. Other chips and memory capacities were not tested. Logical reads include OS-cache hits and do not measure physical SSD traffic; overlapping phase timings are not additive wall-clock totals. No inference speedup is claimed.

The app is arm64 and ad-hoc signed, not notarized. The archive has a SHA-256 checksum and a signed Sparkle update feed. See the repository's model and release validation reports for all repetitions, settings and limitations.

AI assistance: Codex helped implement, review and document these changes. Validation results were collected from the local checks described above.