Repository navigation
v1.0.0
LLM Proxy v1.0.0
An OpenAI-compatible API proxy that sits between your applications and multiple LLM providers (OpenAI, Anthropic, Google, DeepSeek, and others). Route requests through a single endpoint and let the proxy handle the rest.
What it does
- Multi-provider routing — Define any number of providers and models in a single config file. One API endpoint, multiple backends.
- Automatic failover — If a provider fails, requests automatically fall through to the next available one. Circuit breakers trip after 3 consecutive failures and recover automatically.
- Background probing & auto-switch — Continuously tests provider latency and throughput, then switches to the best-performing provider automatically.
- Config hot-reload — Edit
config.yamland changes take effect within seconds. No restart needed. Also supportskill -USR1 <pid>for on-demand reload. Invalid config is rejected gracefully. - Streaming-first — SSE streaming by default, with full token usage tracking and per-provider TTFT/TPS metrics.
- OpenAI-compatible — Drop-in replacement for the OpenAI API. Works with any client that targets
/v1/chat/completionsor/v1/models. - Rate limit handling — Respects
429 Retry-Afterheaders with non-blocking retries. - Optional auth — Protect your proxy with a bearer token so only authorised clients can access it.
- Prometheus metrics — Request counts, durations, and per-provider success/failure counters exposed at
/metrics. - Docker-ready — Ships with a
Dockerfileanddocker-compose.yml. One command to get running.