Skip to content

Task 22 evidence v9

Latest

Choose a tag to compare

@MaybeIcanShow MaybeIcanShow released this 11 Aug 16:43

Full Task 22 hybrid async text-generation performance evidence.

  • Tested a frozen dirty workspace snapshot based on commit 4603e03ffde5769813346b313ce7284c8f726640; the complete tested snapshot and provenance metadata are attached.
  • Optimized vs zero-KL response throughput: +12.90% (paired trials: +13.20%, +12.60%, +12.91%).
  • Formal benchmark completion: 9/9 jobs successful, each with 20 steps and 640 samples.
  • Optimized vs baseline response throughput: +13.51%; end-to-end time vs zero-KL: -8.26%.
  • One initial excluded attempt failed during Ray runtime-environment packaging. It was an infrastructure failure and is not included in the formal measurements.
  • raw-evidence.tar.gz SHA-256: e76f53da3ff064da4045b9b539c9aebddeff5a959884ebb884092c1b39106d81.

See report.md, source-provenance.txt, and raw-evidence-index.csv for methodology and reproducibility details.