Skip to content

v2.6.3

Choose a tag to compare

@vitobotta vitobotta released this 01 Jul 13:35
· 11 commits to main since this release
0630a5a

What's Fixed

Inaccurate TPS from short generations no longer skews provider scoring. When a request generated only a few tokens (5–20) and the provider did not report server-side timing, the per-request total_tps was computed from the chunk arrival window — the time between the first and last chunk reaching the proxy. For 5 tokens this might be 0.02s (giving 250 tps) or 0.0s (giving nil). The value was pure network jitter, not decode throughput, and a single such sample could inflate the provider scorer's average enough to trigger an incorrect auto-switch.

build_stream_result now suppresses the arrival-window TPS fallback when completion_tokens is below MIN_ARRIVAL_TPS_TOKENS (50) — the same threshold already used by PERCENTILE_MIN_TOKENS to exclude short samples from percentile computation. Server-side timing (tokens_per_second, completion_time, generation-duration, vLLM server_duration) is always used regardless of token count. When no server timing is available and the generation is short, total_tps is nil and the sample is stored without a TPS value; its token counts still contribute to the rolling aggregate.

record_metrics no longer falls back to content_tps when total_tps is nil for short generations. The previous total_tps || content_tps fallback would leak the equally noisy content-arrival TPS to the scorer, defeating the suppression. The fallback is now gated on completion_tokens >= MIN_ARRIVAL_TPS_TOKENS, so it only applies when the generation is long enough for the arrival-window estimate to be meaningful.

Performance

  • average_metrics (used by score_from_avg for auto-switch decisions) now uses a token-weighted aggregate — sum(tps × tokens) / sum(tokens) — instead of an arithmetic mean. A 5-token request at 250 tps contributes proportionally less than a 500-token request at 60 tps. Falls back to the arithmetic mean when no samples have token counts (providers that never report completion_tokens). This is the same weighting already used by rolling_tps for the periodic TPS log.

Tests

  • 6 new tests: 3 for build_stream_result threshold behaviour (short generation nil TPS, short generation with server timing, long generation arrival-window TPS), 2 for average_metrics token-weighting (weighted aggregate, arithmetic mean fallback), 1 for record_metrics content_tps suppression for short generations.
  • 249 runs, 611 assertions, 0 failures across 5 test files.