Repository navigation
v2.6.3
What's Fixed
Inaccurate TPS from short generations no longer skews provider scoring. When a request generated only a few tokens (5–20) and the provider did not report server-side timing, the per-request total_tps was computed from the chunk arrival window — the time between the first and last chunk reaching the proxy. For 5 tokens this might be 0.02s (giving 250 tps) or 0.0s (giving nil). The value was pure network jitter, not decode throughput, and a single such sample could inflate the provider scorer's average enough to trigger an incorrect auto-switch.
build_stream_result now suppresses the arrival-window TPS fallback when completion_tokens is below MIN_ARRIVAL_TPS_TOKENS (50) — the same threshold already used by PERCENTILE_MIN_TOKENS to exclude short samples from percentile computation. Server-side timing (tokens_per_second, completion_time, generation-duration, vLLM server_duration) is always used regardless of token count. When no server timing is available and the generation is short, total_tps is nil and the sample is stored without a TPS value; its token counts still contribute to the rolling aggregate.
record_metrics no longer falls back to content_tps when total_tps is nil for short generations. The previous total_tps || content_tps fallback would leak the equally noisy content-arrival TPS to the scorer, defeating the suppression. The fallback is now gated on completion_tokens >= MIN_ARRIVAL_TPS_TOKENS, so it only applies when the generation is long enough for the arrival-window estimate to be meaningful.
Performance
average_metrics(used byscore_from_avgfor auto-switch decisions) now uses a token-weighted aggregate —sum(tps × tokens) / sum(tokens)— instead of an arithmetic mean. A 5-token request at 250 tps contributes proportionally less than a 500-token request at 60 tps. Falls back to the arithmetic mean when no samples have token counts (providers that never reportcompletion_tokens). This is the same weighting already used byrolling_tpsfor the periodic TPS log.
Tests
- 6 new tests: 3 for
build_stream_resultthreshold behaviour (short generation nil TPS, short generation with server timing, long generation arrival-window TPS), 2 foraverage_metricstoken-weighting (weighted aggregate, arithmetic mean fallback), 1 forrecord_metricscontent_tps suppression for short generations. - 249 runs, 611 assertions, 0 failures across 5 test files.