Radicalbit AI Gateway is finally here! #89
Pinned
bigmoby
announced in
Announcements
Replies: 2 comments 11 replies
|
🚀 Let's go! |
0 replies
|
Congrats on the launch! How are you handling latency overhead on the guardrails hot-path within the gateway? If you're looking for ultra-fast, sub-millisecond deterministic checks to drop bad inputs before hitting LLM routes, I'd love to see if we can collaborate on a native integration! |
11 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Radicalbit AI Gateway is finally here!
👋 Today we are releasing the first public version of the Radicalbit AI Gateway.
The Gateway sits between your applications and your model providers. Simply point your code at a single endpoint, define your models and routes in a single YAML file, and the Gateway handles the rest: provider credentials, fallbacks, guardrails, caching, and limits. If you are running LLMs in production across more than one provider, this is the layer that keeps your infrastructure from turning into glue code scattered across every service.
📚 Documentation: https://docs.ai-gateway.radicalbit.ai/
What it does
🔀 Multi-strategy routing — keyword, token length, context length, time-of-day (cron), budget ratio, external text classification, and semantic routing via embedding similarity. Cost control becomes a routing decision.
🛡️ Guardrails in the config, not in the code — pattern matching (
contains,starts_with,ends_with,regex), PII detection and redaction via Microsoft Presidio, and LLM-as-a-Judge with your own prompt templates. Enforced at the gateway layer, withblockorwarnbehaviors and per-message-role scoping.⚡ Caching — exact-match and semantic caching with configurable TTL, similarity threshold and distance metric, plus exact caching support for audio transcription requests.
🚦 Limits — rate, token and budget limiting per route, independently combinable, audio-duration-based limiting for transcription routes, and rate limiting for Model Context Protocol (MCP) servers.
🔄 Reliability — automatic model fallback chains that work across providers.
📊 Observability — 20+ Prometheus metrics, end-to-end OpenTelemetry tracing with ClickHouse storage and custom OTLP exporters (Jaeger, Tempo, …). A usage dashboard with cost and token breakdown by route, group, and key. Request tagging and tag-based trace filtering.
🖥️ Built-in UI — projects, configuration editor with load → approve → serve lifecycle, config import/export, groups and API keys, usage dashboard and trace inspection.
🤖 Configure in plain English — describe the route you want and get valid YAML, either from the Generate button in the UI or from your IDE via the Radicalbit Skills plugin.
🚨 Real-time Alerts & Notifications — proactive email notifications triggered on critical gateway events: guardrail interventions (input/output violations, PII detected/redacted, LLM judge breaches). Rules are configurable per route with instant delivery and customizable recipient lists.
🔌 Providers — native support for OpenAI, Anthropic, Google Gemini, DeepSeek, Mistral and Azure OpenAI, plus any OpenAI-compatible endpoint (Ollama, vLLM, OpenRouter, on-premises deployments).
🔌 MCP proxy — expose Model Context Protocol servers through the same route, the same API key with the same trace view as your models. One endpoint per route fans out across several upstreams over Streamable HTTP or stdio, namespacing tools as
{alias}__{tool}.What a route looks like
That is the whole setup. Start the Gateway, send your requests to the
customer-serviceroute, and you get PII filtering, a fallback chain, and a rate limit without a line of application code.Tell us what breaks
We built this based on real production needs, but yours might be different.
All reactions