Skip to content

Default 8k token allowance to allow for reasoning budget.

Latest

Choose a tag to compare

@alexellis alexellis released this 01 Oct 13:45
· 3 commits to master since this release
Allow 8K tokens in default decode samples

Give reasoning-enabled models more room to finish before the decode
completion gate fails. Version the larger allowance as protocol 1.3
without changing prompts, sample counts, prefill, or concurrency.

Keep archived receipts labelled against their original protocol defaults
so the change does not relabel historical benchmark settings.

Signed-off-by: Alex Ellis (OpenFaaS Ltd) <alexellis2@gmail.com>