Repository navigation
Releases: vitobotta/llm-proxy
Release list
v3.3.0
What's New
- Per-model
ttft_timeout. Each model entry can now set its ownttft_timeout, overriding the globaltimeouts.ttftfor streams to that model — useful when one model's reasoning phase legitimately delays the first content token while others must fail fast. Models without the key keep the global default, andttft_timeout: false(or blank) disables the TTFT gate for that model alone. Out-of-range values are rejected at boot and reload, same as the global setting.
Code organisation
ConfigStore.ttft_timeout_for(name)resolves the effective value (model override, elsetimeouts.ttft), andtry_streamtakesmodel_name:to look it up per attempt.config/config.yaml.example, README and AGENTS.md document the new key.
v3.2.0
What's New
-
Custom per-provider request headers. A
headers:mapping on a provider — or on a model's provider entry — is merged into every upstream request to that provider; model-level entries override provider-level ones. Auth and host headers (authorization,host,x-api-key,api-key) are stripped, so a misconfigured header can't override the auth strategy. -
Multiple API keys per provider.
api_keys:takes an ordered list of keys so a single provider entry can drive several subscriptions without duplicating the provider. Each request starts on the first key in order and advances to the next only when the current key hits its rate limit — the same fallback behaviour as providers. A rate-limited key is paused until its reset and skipped; the request falls through to the next provider only once every key is exhausted.api_key:stays the single-key shorthand, andapi_keys:wins when both are set. Per-key pause state is in-memory and resets on restart.
Code organisation
- New
lib/key_rotator.rbowns per-provider key ordering, cooldown, and exhaustion state.config/config.yaml.exampleand AGENTS.md documentheaders:andapi_keys:.
v3.1.1
⚠️ Breaking changes
- Re-tune Prometheus alert thresholds that assumed streamed requests were timed at request start.
request_duration_secondsand in-flight tracking now cover streaming requests end-to-end, so recorded durations for streamed requests will be longer than before.
What's Fixed
- A stream that dies mid-response no longer falls through to the next provider. Previously a client holding part of provider A's stream could receive provider B's response appended after it, producing a garbled double stream. Partial-stream disconnects now stop fallback entirely — across providers and across retry rounds.
kill -USR1 <pid>triggers a config reload again. The signal handler raisedThreadError: can't be called from trap contexton Ruby 3+ and silently stopped working; it now sets a flag the config watcher acts on.- Health and state endpoints no longer fail when a probe has timed out. Failed background probes recorded
InfinityTTFT, which broke JSON serialization of/v1/health/detailand the state file. Failed probes now record a finite penalty sample. - Graceful shutdown no longer stalls for its full drain deadline. The in-flight counter could drop below zero on requests rejected during shutdown, preventing the drain from ever reaching zero.
- Pooled connections are recycled at 300 seconds of age again. Reused connections had their creation time re-stamped, so only idle eviction ever applied and long-lived sockets stayed in the pool indefinitely.
- Responses streams that end after real output no longer receive a spurious
response.failed. A terminal event without a usage block now counts as a proper end-of-stream; only streams that never delivered output get the synthetic failure event. - A bad client request can no longer evict a healthy provider. Client-shape 4xx responses (400/404/422, …) no longer count toward the circuit breaker. 401/403 responses — the provider refusing the proxy's credentials — still do.
- Auto-switch never promotes a circuit-broken provider to active, regardless of how good its recorded samples are.
- Models configured with the same provider name on multiple entries keep independent stats and circuit state. Previously they shared one pool of samples and one circuit.
- A config reload that removes a model's selector returns 503 instead of raising.
- The container boots on case-sensitive filesystems.
require "english"matched the stdlib'sEnglish.rbonly on case-insensitive filesystems, crash-looping the Docker image on Linux.
What's New
- Non-streaming responses relay upstream headers. Content type, request ids (
x-request-id, trace ids), and rate-limit headers now reach the client instead of being replaced by a forcedapplication/json. - Upstream response bodies are read with a hard memory cap. Response bodies are capped at 64 MB and error bodies at 5 MB, streamed rather than buffered whole, so a buggy or hostile upstream can't force unbounded memory use. Oversized success bodies fail the attempt and fall back instead of forwarding truncated JSON.
Code organisation
- AGENTS.md: key-files table completed (
lib/routes/*,lib/tps_reporter.rb,lib/state_persistence.rb),primary:write timing corrected (graceful shutdown, not auto-switch), test commands updated, and gotchas documented for trap-safe signal handlers and lazily-evaluated Sinatra stream blocks.
v3.1.0
What's New
- OpenAI Responses API passthrough on
POST /v1/responses. Every configured model serves the Responses API, streaming and non-streaming, with no opt-in. Requests and SSE events are forwarded to upstream/responsesverbatim; the proxy rewrites onlymodelandstreamand never translates between Responses and chat formats. Chat-only injections (stream_options.include_usage, Fireworksperf_metrics_in_response) are not sent in Responses mode. - Reliable event recognition across provider dialects. Stream parsing understands
event:line event types (OpenAI's SSE layout),output_text.done/reasoning_text.donewithout a preceding delta, and announce-y container events (response.created,response.output_item.added, …) whose embeddedcontent/textfields were previously misread as output tokens. CRLF termination and terminal events split across network chunks are handled via a bounded stream-tail re-parse. - Streams that die without responding now fail fast. A Responses stream that ends with no output deltas — or only container events — receives a synthetic
response.failed/response.not_foundevent withcode: upstream_stoppedbeforedata: [DONE], so clients get a definitive error instead of hanging. The stream ends with exactly one[DONE], regardless of how the provider terminated it. - Errors in Responses mode are Responses-shaped. Internal errors are sent as
response.failedevents withcode: proxy_error. Providers that do not serve/responsesanswer with their own upstream error, which the fallback loop surfaces. - Metric tracking stays chat-completions-only. Responses streams are not token-tracked unless
RESPONSES_ENABLE_THINKING_TRACKING=1is set, which is also the only condition that enables the TTFT gate in Responses mode. Probes always measurechat/completionsregardless of the request's API format.
v3.0.0
⚠️ Breaking changes
- Set
"stream": trueon completion requests that require server-sent events. The/v1/chat/completionsand/v1/completionsendpoints now return one complete JSON response whenstreamis omitted. Explicitstream: falserequests remain non-streaming.
What's New
- Completion streaming is now opt-in, matching the default request behaviour expected by OpenAI-compatible clients.
- The README and contributor documentation now describe both response modes and the explicit streaming opt-in.
Tests
- Route coverage verifies omitted, false, and true
streamvalues. - Docker integration checks verify both complete JSON responses and server-sent event streams.
v2.6.4
What's Fixed
Upstream connection drops are no longer misclassified as client disconnects. When a streaming request failed after chunks had already been forwarded to the client, try_stream caught errors from both sides of the proxy — upstream reads (response.read_body) and client writes (out << chunk) — in a single rescue block. Any IOError, EOFError, ECONNRESET, or Net::ReadTimeout from the upstream provider was converted to ClientDisconnected and logged as client_disconnect, making it appear as though the client had closed the connection when the upstream provider was actually the one that dropped.
Upstream errors that occur after partial stream are now wrapped in a new StreamPartiallySent exception and classified as upstream_disconnect in logs and metrics. Both client_disconnect and upstream_disconnect prevent retries (the client already has partial data), but only upstream_disconnect counts toward the circuit breaker — a client going away is not a provider failure.
Client disconnects no longer penalise providers or cascade through the fallback loop. Previously, when a client disconnected mid-stream, record_failure was called on the active provider, and the request continued to the next fallback provider — which would also fail to write to the dead client socket and also get record_failure. A single client disconnect could open circuits on every provider. with_auto_select now skips record_failure for client_disconnect and breaks out of both the provider loop and the rounds loop immediately.
Tests
- 7 new tests:
StreamPartiallySentdoes not retry and preserves the original error;ClientDisconnecteddoes not retry;failure_reasonreturnsupstream_disconnectfor the new error;with_auto_selectdoes not callrecord_failurefor client disconnects;with_auto_selectbreaks the provider loop on client disconnect;with_auto_selectbreaks the rounds loop on client disconnect;with_auto_selectstill callsrecord_failureforupstream_disconnect. 1 existing test updated to assert the newStreamPartiallySenterror string. - 129 runs, 301 assertions, 0 failures across 9 test files.
v2.6.3
What's Fixed
Inaccurate TPS from short generations no longer skews provider scoring. When a request generated only a few tokens (5–20) and the provider did not report server-side timing, the per-request total_tps was computed from the chunk arrival window — the time between the first and last chunk reaching the proxy. For 5 tokens this might be 0.02s (giving 250 tps) or 0.0s (giving nil). The value was pure network jitter, not decode throughput, and a single such sample could inflate the provider scorer's average enough to trigger an incorrect auto-switch.
build_stream_result now suppresses the arrival-window TPS fallback when completion_tokens is below MIN_ARRIVAL_TPS_TOKENS (50) — the same threshold already used by PERCENTILE_MIN_TOKENS to exclude short samples from percentile computation. Server-side timing (tokens_per_second, completion_time, generation-duration, vLLM server_duration) is always used regardless of token count. When no server timing is available and the generation is short, total_tps is nil and the sample is stored without a TPS value; its token counts still contribute to the rolling aggregate.
record_metrics no longer falls back to content_tps when total_tps is nil for short generations. The previous total_tps || content_tps fallback would leak the equally noisy content-arrival TPS to the scorer, defeating the suppression. The fallback is now gated on completion_tokens >= MIN_ARRIVAL_TPS_TOKENS, so it only applies when the generation is long enough for the arrival-window estimate to be meaningful.
Performance
average_metrics(used byscore_from_avgfor auto-switch decisions) now uses a token-weighted aggregate —sum(tps × tokens) / sum(tokens)— instead of an arithmetic mean. A 5-token request at 250 tps contributes proportionally less than a 500-token request at 60 tps. Falls back to the arithmetic mean when no samples have token counts (providers that never reportcompletion_tokens). This is the same weighting already used byrolling_tpsfor the periodic TPS log.
Tests
- 6 new tests: 3 for
build_stream_resultthreshold behaviour (short generation nil TPS, short generation with server timing, long generation arrival-window TPS), 2 foraverage_metricstoken-weighting (weighted aggregate, arithmetic mean fallback), 1 forrecord_metricscontent_tps suppression for short generations. - 249 runs, 611 assertions, 0 failures across 5 test files.
v2.6.2
What's Fixed
Proactive TTFT timer that fires when a provider sends no data at all. The TTFT timeout check in consume_stream only ran when a chunk arrived — if the provider sent nothing for 31 seconds then sent content, the check never got a chance to fire because first_token was set by the very chunk that triggered the check. The request succeeded with ttft=31.119s despite the 10-second TTFT threshold.
consume_stream now starts a background timer thread that closes the http connection after ttft_timeout seconds if no first token has arrived. This breaks read_body out of its blocking wait. The IOError from the closed socket is caught and converted to TTFTTimeoutError, which retries like any failure. The timer is cancelled as soon as first_token is set.
Stale pooled connections still fail fast with EOFError — the timer only fires on genuinely slow connections (socket alive, no data for ttft_timeout seconds). read_timeout is not lowered; the previous approach of lowering it (commit bf4c545, reverted in 57129b4) broke the EOFError retry path on stale connections.
v2.6.1
What's Fixed
TTFT timeout measured per-attempt instead of from overall request start. The TTFT timeout compared (now - @request_start) against ttft_timeout, where @request_start is set once at the beginning of the overall client request lifecycle (in the before block). By the time a later attempt started — after retries or provider fallbacks — the elapsed time already exceeded ttft_timeout, causing TTFTTimeoutError to fire immediately on the first chunk, even if the new provider would have responded instantly. This produced a rapid loop through all providers and rounds as each one failed on arrival.
The same @request_start was used by build_stream_result to compute the reported TTFT metric, so upstream_ttft_seconds and the logged TTFT were also inflated with prior-attempt time.
try_stream now captures attempt_start at the top of each retry block and passes it to consume_stream and build_stream_result. The TTFT budget is measured from the start of each individual upstream attempt. The reported TTFT metric reflects the actual upstream provider response time for the successful attempt.
v2.6.0
What's New
-
TTFT timeout for streaming requests — a new
timeouts.ttftconfig key (seconds) lets the proxy fall back to the next provider when the primary takes too long to start generating. When a streaming request doesn't receive a first token (thinking or content) within the configured deadline, the proxy treats it as a failure: it retries on the same provider up tomax_attempts, then falls back to the next provider — the same path as any other failure.timeouts: open: 30 read: 300 write: 60 ttft: 15 # max seconds to wait for the first token (streaming only)
Set
ttfttonullor omit it to disable the feature.How it works: An application-level check in
consume_streamuses a total deadline (request_start + ttft_timeout) to detect when no first token has arrived. The check runs after each parsed chunk — iftracker.first_tokenis still nil and the deadline has passed,TTFTTimeoutErroris raised before the chunk is forwarded to the client.Stale pooled connections are not affected — they continue to be caught by the existing
EOFErrorretry path (<1s, free, doesn't consume the attempt budget). The normalread_timeout(300s) handles truly dead connections.The
TTFTTimeoutErrorbypasses the mid-stream corruption guard (which normally prevents retries after data has been forwarded) — safe because no content was sent to the client, only comment lines that SSE clients ignore.Scope: streaming only. Non-streaming requests and background probes are unaffected. Requires
tracking.enabled: true(the default) — chunk parsing is needed to detect the first token.
Performance
HTTPSupport::TTFTTimeoutError— new exception class, rescued with a distinct"TTFT timeout"label in logs and"ttft_timeout"in Prometheus metrics.
Configuration
timeouts.ttft— new optional key. Validated as a positive number (1..86400 seconds). When omitted, the feature is disabled and behaviour is unchanged.
Tests
- 15 new tests: 6 for
consume_streamTTFT behaviour, 4 fortry_stream(timeout detection, content within deadline, ping within deadline, tracking-disabled skip), 5 for config validation. - 59 runs, 130 assertions, 0 failures.