Skip to content

v3.1.0

Choose a tag to compare

@vitobotta vitobotta released this 19 Aug 20:01
· 7 commits to main since this release
91ae11a

What's New

  • OpenAI Responses API passthrough on POST /v1/responses. Every configured model serves the Responses API, streaming and non-streaming, with no opt-in. Requests and SSE events are forwarded to upstream /responses verbatim; the proxy rewrites only model and stream and never translates between Responses and chat formats. Chat-only injections (stream_options.include_usage, Fireworks perf_metrics_in_response) are not sent in Responses mode.
  • Reliable event recognition across provider dialects. Stream parsing understands event: line event types (OpenAI's SSE layout), output_text.done / reasoning_text.done without a preceding delta, and announce-y container events (response.created, response.output_item.added, …) whose embedded content/text fields were previously misread as output tokens. CRLF termination and terminal events split across network chunks are handled via a bounded stream-tail re-parse.
  • Streams that die without responding now fail fast. A Responses stream that ends with no output deltas — or only container events — receives a synthetic response.failed / response.not_found event with code: upstream_stopped before data: [DONE], so clients get a definitive error instead of hanging. The stream ends with exactly one [DONE], regardless of how the provider terminated it.
  • Errors in Responses mode are Responses-shaped. Internal errors are sent as response.failed events with code: proxy_error. Providers that do not serve /responses answer with their own upstream error, which the fallback loop surfaces.
  • Metric tracking stays chat-completions-only. Responses streams are not token-tracked unless RESPONSES_ENABLE_THINKING_TRACKING=1 is set, which is also the only condition that enables the TTFT gate in Responses mode. Probes always measure chat/completions regardless of the request's API format.