Skip to content

v3.1.1

Choose a tag to compare

@vitobotta vitobotta released this 04 Oct 19:54
· 3 commits to main since this release
8181e87

⚠️ Breaking changes

  • Re-tune Prometheus alert thresholds that assumed streamed requests were timed at request start. request_duration_seconds and in-flight tracking now cover streaming requests end-to-end, so recorded durations for streamed requests will be longer than before.

What's Fixed

  • A stream that dies mid-response no longer falls through to the next provider. Previously a client holding part of provider A's stream could receive provider B's response appended after it, producing a garbled double stream. Partial-stream disconnects now stop fallback entirely — across providers and across retry rounds.
  • kill -USR1 <pid> triggers a config reload again. The signal handler raised ThreadError: can't be called from trap context on Ruby 3+ and silently stopped working; it now sets a flag the config watcher acts on.
  • Health and state endpoints no longer fail when a probe has timed out. Failed background probes recorded Infinity TTFT, which broke JSON serialization of /v1/health/detail and the state file. Failed probes now record a finite penalty sample.
  • Graceful shutdown no longer stalls for its full drain deadline. The in-flight counter could drop below zero on requests rejected during shutdown, preventing the drain from ever reaching zero.
  • Pooled connections are recycled at 300 seconds of age again. Reused connections had their creation time re-stamped, so only idle eviction ever applied and long-lived sockets stayed in the pool indefinitely.
  • Responses streams that end after real output no longer receive a spurious response.failed. A terminal event without a usage block now counts as a proper end-of-stream; only streams that never delivered output get the synthetic failure event.
  • A bad client request can no longer evict a healthy provider. Client-shape 4xx responses (400/404/422, …) no longer count toward the circuit breaker. 401/403 responses — the provider refusing the proxy's credentials — still do.
  • Auto-switch never promotes a circuit-broken provider to active, regardless of how good its recorded samples are.
  • Models configured with the same provider name on multiple entries keep independent stats and circuit state. Previously they shared one pool of samples and one circuit.
  • A config reload that removes a model's selector returns 503 instead of raising.
  • The container boots on case-sensitive filesystems. require "english" matched the stdlib's English.rb only on case-insensitive filesystems, crash-looping the Docker image on Linux.

What's New

  • Non-streaming responses relay upstream headers. Content type, request ids (x-request-id, trace ids), and rate-limit headers now reach the client instead of being replaced by a forced application/json.
  • Upstream response bodies are read with a hard memory cap. Response bodies are capped at 64 MB and error bodies at 5 MB, streamed rather than buffered whole, so a buggy or hostile upstream can't force unbounded memory use. Oversized success bodies fail the attempt and fall back instead of forwarding truncated JSON.

Code organisation

  • AGENTS.md: key-files table completed (lib/routes/*, lib/tps_reporter.rb, lib/state_persistence.rb), primary: write timing corrected (graceful shutdown, not auto-switch), test commands updated, and gotchas documented for trap-safe signal handlers and lazily-evaluated Sinatra stream blocks.