Skip to content

v1.0.0

Choose a tag to compare

@vitobotta vitobotta released this 09 May 21:11
· 61 commits to main since this release
9a53897

LLM Proxy v1.0.0

An OpenAI-compatible API proxy that sits between your applications and multiple LLM providers (OpenAI, Anthropic, Google, DeepSeek, and others). Route requests through a single endpoint and let the proxy handle the rest.

What it does

  • Multi-provider routing — Define any number of providers and models in a single config file. One API endpoint, multiple backends.
  • Automatic failover — If a provider fails, requests automatically fall through to the next available one. Circuit breakers trip after 3 consecutive failures and recover automatically.
  • Background probing & auto-switch — Continuously tests provider latency and throughput, then switches to the best-performing provider automatically.
  • Config hot-reload — Edit config.yaml and changes take effect within seconds. No restart needed. Also supports kill -USR1 <pid> for on-demand reload. Invalid config is rejected gracefully.
  • Streaming-first — SSE streaming by default, with full token usage tracking and per-provider TTFT/TPS metrics.
  • OpenAI-compatible — Drop-in replacement for the OpenAI API. Works with any client that targets /v1/chat/completions or /v1/models.
  • Rate limit handling — Respects 429 Retry-After headers with non-blocking retries.
  • Optional auth — Protect your proxy with a bearer token so only authorised clients can access it.
  • Prometheus metrics — Request counts, durations, and per-provider success/failure counters exposed at /metrics.
  • Docker-ready — Ships with a Dockerfile and docker-compose.yml. One command to get running.