Repository navigation
v2.6.1
What's Fixed
TTFT timeout measured per-attempt instead of from overall request start. The TTFT timeout compared (now - @request_start) against ttft_timeout, where @request_start is set once at the beginning of the overall client request lifecycle (in the before block). By the time a later attempt started — after retries or provider fallbacks — the elapsed time already exceeded ttft_timeout, causing TTFTTimeoutError to fire immediately on the first chunk, even if the new provider would have responded instantly. This produced a rapid loop through all providers and rounds as each one failed on arrival.
The same @request_start was used by build_stream_result to compute the reported TTFT metric, so upstream_ttft_seconds and the logged TTFT were also inflated with prior-attempt time.
try_stream now captures attempt_start at the top of each retry block and passes it to consume_stream and build_stream_result. The TTFT budget is measured from the start of each individual upstream attempt. The reported TTFT metric reflects the actual upstream provider response time for the successful attempt.