Skip to content

Releases: vitobotta/llm-proxy

v3.3.0

Choose a tag to compare

@vitobotta vitobotta released this 07 Oct 21:12
dd8dced

What's New

  • Per-model ttft_timeout. Each model entry can now set its own ttft_timeout, overriding the global timeouts.ttft for streams to that model — useful when one model's reasoning phase legitimately delays the first content token while others must fail fast. Models without the key keep the global default, and ttft_timeout: false (or blank) disables the TTFT gate for that model alone. Out-of-range values are rejected at boot and reload, same as the global setting.

Code organisation

  • ConfigStore.ttft_timeout_for(name) resolves the effective value (model override, else timeouts.ttft), and try_stream takes model_name: to look it up per attempt. config/config.yaml.example, README and AGENTS.md document the new key.

v3.2.0

Choose a tag to compare

@vitobotta vitobotta released this 05 Oct 19:56
e92ff70

What's New

  • Custom per-provider request headers. A headers: mapping on a provider — or on a model's provider entry — is merged into every upstream request to that provider; model-level entries override provider-level ones. Auth and host headers (authorization, host, x-api-key, api-key) are stripped, so a misconfigured header can't override the auth strategy.

  • Multiple API keys per provider. api_keys: takes an ordered list of keys so a single provider entry can drive several subscriptions without duplicating the provider. Each request starts on the first key in order and advances to the next only when the current key hits its rate limit — the same fallback behaviour as providers. A rate-limited key is paused until its reset and skipped; the request falls through to the next provider only once every key is exhausted. api_key: stays the single-key shorthand, and api_keys: wins when both are set. Per-key pause state is in-memory and resets on restart.

Code organisation

  • New lib/key_rotator.rb owns per-provider key ordering, cooldown, and exhaustion state. config/config.yaml.example and AGENTS.md document headers: and api_keys:.

v3.1.1

Choose a tag to compare

@vitobotta vitobotta released this 04 Oct 19:54
8181e87

⚠️ Breaking changes

  • Re-tune Prometheus alert thresholds that assumed streamed requests were timed at request start. request_duration_seconds and in-flight tracking now cover streaming requests end-to-end, so recorded durations for streamed requests will be longer than before.

What's Fixed

  • A stream that dies mid-response no longer falls through to the next provider. Previously a client holding part of provider A's stream could receive provider B's response appended after it, producing a garbled double stream. Partial-stream disconnects now stop fallback entirely — across providers and across retry rounds.
  • kill -USR1 <pid> triggers a config reload again. The signal handler raised ThreadError: can't be called from trap context on Ruby 3+ and silently stopped working; it now sets a flag the config watcher acts on.
  • Health and state endpoints no longer fail when a probe has timed out. Failed background probes recorded Infinity TTFT, which broke JSON serialization of /v1/health/detail and the state file. Failed probes now record a finite penalty sample.
  • Graceful shutdown no longer stalls for its full drain deadline. The in-flight counter could drop below zero on requests rejected during shutdown, preventing the drain from ever reaching zero.
  • Pooled connections are recycled at 300 seconds of age again. Reused connections had their creation time re-stamped, so only idle eviction ever applied and long-lived sockets stayed in the pool indefinitely.
  • Responses streams that end after real output no longer receive a spurious response.failed. A terminal event without a usage block now counts as a proper end-of-stream; only streams that never delivered output get the synthetic failure event.
  • A bad client request can no longer evict a healthy provider. Client-shape 4xx responses (400/404/422, …) no longer count toward the circuit breaker. 401/403 responses — the provider refusing the proxy's credentials — still do.
  • Auto-switch never promotes a circuit-broken provider to active, regardless of how good its recorded samples are.
  • Models configured with the same provider name on multiple entries keep independent stats and circuit state. Previously they shared one pool of samples and one circuit.
  • A config reload that removes a model's selector returns 503 instead of raising.
  • The container boots on case-sensitive filesystems. require "english" matched the stdlib's English.rb only on case-insensitive filesystems, crash-looping the Docker image on Linux.

What's New

  • Non-streaming responses relay upstream headers. Content type, request ids (x-request-id, trace ids), and rate-limit headers now reach the client instead of being replaced by a forced application/json.
  • Upstream response bodies are read with a hard memory cap. Response bodies are capped at 64 MB and error bodies at 5 MB, streamed rather than buffered whole, so a buggy or hostile upstream can't force unbounded memory use. Oversized success bodies fail the attempt and fall back instead of forwarding truncated JSON.

Code organisation

  • AGENTS.md: key-files table completed (lib/routes/*, lib/tps_reporter.rb, lib/state_persistence.rb), primary: write timing corrected (graceful shutdown, not auto-switch), test commands updated, and gotchas documented for trap-safe signal handlers and lazily-evaluated Sinatra stream blocks.

v3.1.0

Choose a tag to compare

@vitobotta vitobotta released this 19 Aug 20:01
91ae11a

What's New

  • OpenAI Responses API passthrough on POST /v1/responses. Every configured model serves the Responses API, streaming and non-streaming, with no opt-in. Requests and SSE events are forwarded to upstream /responses verbatim; the proxy rewrites only model and stream and never translates between Responses and chat formats. Chat-only injections (stream_options.include_usage, Fireworks perf_metrics_in_response) are not sent in Responses mode.
  • Reliable event recognition across provider dialects. Stream parsing understands event: line event types (OpenAI's SSE layout), output_text.done / reasoning_text.done without a preceding delta, and announce-y container events (response.created, response.output_item.added, …) whose embedded content/text fields were previously misread as output tokens. CRLF termination and terminal events split across network chunks are handled via a bounded stream-tail re-parse.
  • Streams that die without responding now fail fast. A Responses stream that ends with no output deltas — or only container events — receives a synthetic response.failed / response.not_found event with code: upstream_stopped before data: [DONE], so clients get a definitive error instead of hanging. The stream ends with exactly one [DONE], regardless of how the provider terminated it.
  • Errors in Responses mode are Responses-shaped. Internal errors are sent as response.failed events with code: proxy_error. Providers that do not serve /responses answer with their own upstream error, which the fallback loop surfaces.
  • Metric tracking stays chat-completions-only. Responses streams are not token-tracked unless RESPONSES_ENABLE_THINKING_TRACKING=1 is set, which is also the only condition that enables the TTFT gate in Responses mode. Probes always measure chat/completions regardless of the request's API format.

v3.0.0

Choose a tag to compare

@vitobotta vitobotta released this 09 Aug 19:45
03a9107

⚠️ Breaking changes

  • Set "stream": true on completion requests that require server-sent events. The /v1/chat/completions and /v1/completions endpoints now return one complete JSON response when stream is omitted. Explicit stream: false requests remain non-streaming.

What's New

  • Completion streaming is now opt-in, matching the default request behaviour expected by OpenAI-compatible clients.
  • The README and contributor documentation now describe both response modes and the explicit streaming opt-in.

Tests

  • Route coverage verifies omitted, false, and true stream values.
  • Docker integration checks verify both complete JSON responses and server-sent event streams.

v2.6.4

Choose a tag to compare

@vitobotta vitobotta released this 07 Jul 19:27
cec70cc

What's Fixed

Upstream connection drops are no longer misclassified as client disconnects. When a streaming request failed after chunks had already been forwarded to the client, try_stream caught errors from both sides of the proxy — upstream reads (response.read_body) and client writes (out << chunk) — in a single rescue block. Any IOError, EOFError, ECONNRESET, or Net::ReadTimeout from the upstream provider was converted to ClientDisconnected and logged as client_disconnect, making it appear as though the client had closed the connection when the upstream provider was actually the one that dropped.

Upstream errors that occur after partial stream are now wrapped in a new StreamPartiallySent exception and classified as upstream_disconnect in logs and metrics. Both client_disconnect and upstream_disconnect prevent retries (the client already has partial data), but only upstream_disconnect counts toward the circuit breaker — a client going away is not a provider failure.

Client disconnects no longer penalise providers or cascade through the fallback loop. Previously, when a client disconnected mid-stream, record_failure was called on the active provider, and the request continued to the next fallback provider — which would also fail to write to the dead client socket and also get record_failure. A single client disconnect could open circuits on every provider. with_auto_select now skips record_failure for client_disconnect and breaks out of both the provider loop and the rounds loop immediately.

Tests

  • 7 new tests: StreamPartiallySent does not retry and preserves the original error; ClientDisconnected does not retry; failure_reason returns upstream_disconnect for the new error; with_auto_select does not call record_failure for client disconnects; with_auto_select breaks the provider loop on client disconnect; with_auto_select breaks the rounds loop on client disconnect; with_auto_select still calls record_failure for upstream_disconnect. 1 existing test updated to assert the new StreamPartiallySent error string.
  • 129 runs, 301 assertions, 0 failures across 9 test files.

v2.6.3

Choose a tag to compare

@vitobotta vitobotta released this 01 Jul 13:35
0630a5a

What's Fixed

Inaccurate TPS from short generations no longer skews provider scoring. When a request generated only a few tokens (5–20) and the provider did not report server-side timing, the per-request total_tps was computed from the chunk arrival window — the time between the first and last chunk reaching the proxy. For 5 tokens this might be 0.02s (giving 250 tps) or 0.0s (giving nil). The value was pure network jitter, not decode throughput, and a single such sample could inflate the provider scorer's average enough to trigger an incorrect auto-switch.

build_stream_result now suppresses the arrival-window TPS fallback when completion_tokens is below MIN_ARRIVAL_TPS_TOKENS (50) — the same threshold already used by PERCENTILE_MIN_TOKENS to exclude short samples from percentile computation. Server-side timing (tokens_per_second, completion_time, generation-duration, vLLM server_duration) is always used regardless of token count. When no server timing is available and the generation is short, total_tps is nil and the sample is stored without a TPS value; its token counts still contribute to the rolling aggregate.

record_metrics no longer falls back to content_tps when total_tps is nil for short generations. The previous total_tps || content_tps fallback would leak the equally noisy content-arrival TPS to the scorer, defeating the suppression. The fallback is now gated on completion_tokens >= MIN_ARRIVAL_TPS_TOKENS, so it only applies when the generation is long enough for the arrival-window estimate to be meaningful.

Performance

  • average_metrics (used by score_from_avg for auto-switch decisions) now uses a token-weighted aggregate — sum(tps × tokens) / sum(tokens) — instead of an arithmetic mean. A 5-token request at 250 tps contributes proportionally less than a 500-token request at 60 tps. Falls back to the arithmetic mean when no samples have token counts (providers that never report completion_tokens). This is the same weighting already used by rolling_tps for the periodic TPS log.

Tests

  • 6 new tests: 3 for build_stream_result threshold behaviour (short generation nil TPS, short generation with server timing, long generation arrival-window TPS), 2 for average_metrics token-weighting (weighted aggregate, arithmetic mean fallback), 1 for record_metrics content_tps suppression for short generations.
  • 249 runs, 611 assertions, 0 failures across 5 test files.

v2.6.2

Choose a tag to compare

@vitobotta vitobotta released this 01 Jul 11:22
e09435c

What's Fixed

Proactive TTFT timer that fires when a provider sends no data at all. The TTFT timeout check in consume_stream only ran when a chunk arrived — if the provider sent nothing for 31 seconds then sent content, the check never got a chance to fire because first_token was set by the very chunk that triggered the check. The request succeeded with ttft=31.119s despite the 10-second TTFT threshold.

consume_stream now starts a background timer thread that closes the http connection after ttft_timeout seconds if no first token has arrived. This breaks read_body out of its blocking wait. The IOError from the closed socket is caught and converted to TTFTTimeoutError, which retries like any failure. The timer is cancelled as soon as first_token is set.

Stale pooled connections still fail fast with EOFError — the timer only fires on genuinely slow connections (socket alive, no data for ttft_timeout seconds). read_timeout is not lowered; the previous approach of lowering it (commit bf4c545, reverted in 57129b4) broke the EOFError retry path on stale connections.

v2.6.1

Choose a tag to compare

@vitobotta vitobotta released this 01 Jul 10:52
a1ebe24

What's Fixed

TTFT timeout measured per-attempt instead of from overall request start. The TTFT timeout compared (now - @request_start) against ttft_timeout, where @request_start is set once at the beginning of the overall client request lifecycle (in the before block). By the time a later attempt started — after retries or provider fallbacks — the elapsed time already exceeded ttft_timeout, causing TTFTTimeoutError to fire immediately on the first chunk, even if the new provider would have responded instantly. This produced a rapid loop through all providers and rounds as each one failed on arrival.

The same @request_start was used by build_stream_result to compute the reported TTFT metric, so upstream_ttft_seconds and the logged TTFT were also inflated with prior-attempt time.

try_stream now captures attempt_start at the top of each retry block and passes it to consume_stream and build_stream_result. The TTFT budget is measured from the start of each individual upstream attempt. The reported TTFT metric reflects the actual upstream provider response time for the successful attempt.

v2.6.0

Choose a tag to compare

@vitobotta vitobotta released this 30 Jun 19:22
bf4c545

What's New

  • TTFT timeout for streaming requests — a new timeouts.ttft config key (seconds) lets the proxy fall back to the next provider when the primary takes too long to start generating. When a streaming request doesn't receive a first token (thinking or content) within the configured deadline, the proxy treats it as a failure: it retries on the same provider up to max_attempts, then falls back to the next provider — the same path as any other failure.

    timeouts:
      open: 30
      read: 300
      write: 60
      ttft: 15   # max seconds to wait for the first token (streaming only)

    Set ttft to null or omit it to disable the feature.

    How it works: An application-level check in consume_stream uses a total deadline (request_start + ttft_timeout) to detect when no first token has arrived. The check runs after each parsed chunk — if tracker.first_token is still nil and the deadline has passed, TTFTTimeoutError is raised before the chunk is forwarded to the client.

    Stale pooled connections are not affected — they continue to be caught by the existing EOFError retry path (<1s, free, doesn't consume the attempt budget). The normal read_timeout (300s) handles truly dead connections.

    The TTFTTimeoutError bypasses the mid-stream corruption guard (which normally prevents retries after data has been forwarded) — safe because no content was sent to the client, only comment lines that SSE clients ignore.

    Scope: streaming only. Non-streaming requests and background probes are unaffected. Requires tracking.enabled: true (the default) — chunk parsing is needed to detect the first token.

Performance

  • HTTPSupport::TTFTTimeoutError — new exception class, rescued with a distinct "TTFT timeout" label in logs and "ttft_timeout" in Prometheus metrics.

Configuration

  • timeouts.ttft — new optional key. Validated as a positive number (1..86400 seconds). When omitted, the feature is disabled and behaviour is unchanged.

Tests

  • 15 new tests: 6 for consume_stream TTFT behaviour, 4 for try_stream (timeout detection, content within deadline, ping within deadline, tracking-disabled skip), 5 for config validation.
  • 59 runs, 130 assertions, 0 failures.