Skip to content

History / Benchmarking

Revisions

  • Add GPT-OSS 120B benchmark results and MoE troubleshooting Benchmarking: Added MoE test results (100% acceptance, 0.39x speedup) and target speed threshold table showing when speculation helps vs hurts. Troubleshooting: Added MoE RPC distribution failure section and version matching doctor check note. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

    @youngharold youngharold committed Feb 20, 2026
  • docs: add cross-family benchmark results and families script docs Add Llama 70B, Qwen3 235B, Qwen3.5 397B results to Benchmarking page. New summary table comparing all cloud configs. Document benchmark_families.py usage and config system. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

    @youngharold youngharold committed Feb 18, 2026
  • docs: add Benchmarking page, update Home and Speculative Decoding New Benchmarking wiki page with full reproduction instructions for both benchmark scripts, methodology notes, and published results table. Add Llama 3.1 8B → 405B OpenRouter results to Speculative Decoding page. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

    @youngharold youngharold committed Feb 18, 2026