Releases: diegosouzapw/OmniRoute
Release list
v3.8.51
OmniRoute v3.8.51
2,022 documented changes — 235 features, 1,454 fixes and 333 maintenance entries from 1,973 commits in the cycle, with 301 external contributors. Thank you all.
📄 Complete release notes (every entry, full description and credit): attached below as CHANGELOG-v3.8.51.md, and in CHANGELOG.md → [3.8.51]. GitHub caps a release body at 125,000 characters, so the 1,454 fixes and 333 maintenance entries live there in full.
📊 Release by the numbers
| 👥 People who contributed | 325 |
| 📝 Commits in the cycle | 1,972 |
| 🔀 Pull requests referenced | 1,942 |
| 📋 Changelog entries | 2,022 |
| 🙌 Contributors credited in entries | 322 |
| 🤖 Automated dependency commits | 33 |
Entries by type
| Type | Count |
|---|---|
| 🐛 Fixes | 1443 |
| ✨ Features | 228 |
| 🧹 Chore | 135 |
| 🧪 Tests | 90 |
| 📚 Docs | 72 |
| ⚙️ CI | 11 |
| ♻️ Refactor | 11 |
| ⚡ Performance | 10 |
| 🏗️ Build | 8 |
| 🔒 Security | 7 |
| 📦 Dependencies | 5 |
| ⏪ Reverts | 2 |
🏆 Top 25 contributors this cycle
By commits in 091589089c..4aed4b4a08, author identities consolidated via .mailmap and the merged PR's GitHub login. Bots excluded.
| # | Contributor | Commits |
|---|---|---|
| 🥇 | diegosouzapw | 589 |
| 🥈 | Dizzle (@maxmad64bis) | 175 |
| 🥉 | Bob.Hou (@HouMinXi) | 167 |
| 4 | Paco Cartones (@pacocartones) | 93 |
| 5 | Koosha Paridehpour (@KooshaPari) | 50 |
| 6 | Ravi Tharuma (@RaviTharuma) | 43 |
| 7 | Markus Hartung (@hartmark) | 33 |
| 8 | Nguyen Thanh Dat (@ntdatt812) | 31 |
| 9 | QuangBlue | 31 |
| 10 | Webman (@jonlwheat2-gif) | 23 |
| 11 | Goni Sulaiman (@gonisulaimann) | 22 |
| 12 | lorenzozane (@lorenzozanee) | 22 |
| 13 | anhtahaylove | 19 |
| 14 | backryun | 18 |
| 15 | Fouad Salkini (@fouadSalkini) | 16 |
| 16 | Abhishek Sharma (@abhisheksharma2411) | 15 |
| 17 | Nguyen Thanh Dat (@datrixlab) | 15 |
| 18 | Patryk Kopyciński (@patrykkopycinski) | 15 |
| 19 | Rafa Martins (@rafacpti23) | 15 |
| 20 | Caio Figueiroa (@shipsfromrio) | 15 |
| 21 | initguru | 12 |
| 22 | Jan Leon (@JxnLexn) | 12 |
| 23 | Syed Raheemuddin (@raheemuddin786) | 12 |
| 24 | Xmon Dai (@xiechimon) | 12 |
| 25 | Nguyễn Viết Tuấn (@TheDemonTuan) | 11 |
Highlights
- Providers and catalogs (42 provider features): new gateways and web-session providers (UC / UC Direct, Lyceum, regolo /v1), GPT-6 Astra/Sol/Luna in the Codex and OpenAI catalogs with effort aliases, account-live listings as the chat catalog source for Claude/Codex/Copilot/AGY, safe Codex model discovery.
- Routing and resilience: hierarchical concurrency admission, adaptive reasoning effort (
auto),expiry-firstaccount fallback, per-kind breaker cooldown escalation, provider-breaker close on a successful probe, local-cooldown 429s no longer read as quota exhaustion. - Proxies: pools stop re-serving a member the provider just refused, refused-egress set-aside tuning, egress-IP visibility per pool, selector steering away from recently refused members, persisted blocked-verdict history.
- Video / Audio Bridge: hardened drill-down cache substrate, bounded transcript contract, opt-in audio extraction + STT, earned "embedded" transcript provenance.
- Free tier and Radar: eligibility-gated free-tier bucket,
customModels[].isFree, free-tier capability in the plugin manifest, Radar rate-limit and rules surfacing. - Dashboard and observability: continuous call-log export, orchestration "Repeat", managed-lease sessions view, per-attempt upstream timing on proxy logs (opt-in), logs-export truncation notice.
- Security hardening (late in the cycle): pre-request hooks in an isolated realm (GHSA-9p9m), IP allow/deny judged on the connection, trailing-dot internal hosts, management auth on OAuth routes, OIDC/Trae/GHE Copilot guards, constant-time token comparison, loopback-only version manager, traversal refused in the Codex
/v1/responsessubpath, linear-time tool-markup scanners. - Release engineering: npm Trusted Publishing (OIDC), lockfile valid for npm 10/11/12 again, AI-attribution gate with a reviewed historical allowlist.
✨ New Features
- feat(audio): proxy native ElevenLabs voices, text-to-speech, and speech-to-text HTTP routes through stored OmniRoute credentials, preserving query strings, multipart uploads, binary responses, and upstream errors (#10556 #11434) — thanks @hartmark
- Added Google AI Studio Gemini batch text-to-speech support through
POST /v1/audio/speech(#11434) — thanks @hartmark - Run synchronous RTK and Caveman request compression in a bounded worker-thread pool, keeping large
/v1/responsescompression heaps outside the HTTP isolate while preserving strict fail-open behavior and per-engine telemetry (#11434) — thanks @hartmark - feat(routing): subscription-first auto groupings —
auto/subscriptionroutes only through plan-included connections with a documented hard-stop overage and fails closed on exhaustion, whileauto/thriftyorders the poolsubscription → keyless → free → cheap → premiumand steps up one rung at a time as each is exhausted. Billing class comes from a curated per-connection catalog (uncurated is treated as metered, never plan-included), both reuse STRICT_ZERO_COST's per-connection verification, and a quota reading whoseresetAthas passed is now refreshed regardless of TTL so routing returns to plan… (#11146) — thanks @yourspraveen - feat(providers): publish a management-authenticated versioned web-session credential contract from OmniRoute's canonical browser credential metadata (#11340) — thanks @Zartharas
- feat(video bridge): harden the optional drill-down cache substrate with exact-path broker policy, canonical principal/session/media isolation, independent retained-byte quotas, cancellation-safe commits, rejection of excess or non-canonical Base64 padding and non-JPEG/truncated media, warning-sensitive full JPEG canonicalization that strips trailing polyglot bytes, server-derived dimensions, and auditable derivation metadata; production tenant binding and multi-resolution selection remain follow-up work (#11369 #11434) — thanks @hartmark
- feat(search): Add Xquik X search with typed results, credential validation, REST routing, and MCP selection (#11370) — thanks @kriptoburak
- feat(video): add an opt-in focused analysis mode that safely uses a normalized, 500-code-point latest-user hint for task-aware frame captions while preserving full-mode prompts, temporal-window isolation, and cache identity without storing raw task text (#11383 #11434) — thanks @hartmark
- feat(dashboard): surface durable exclusive managed leases in the existing Sessions view, keeping leased clients visible across idle gaps while marking connections with in-flight work as active (#11389) — thanks @KaspaPulse
- build(bun): allow Turbopack bundler flag on Bun 1.4+ with configurable Webpack fallback (#11471) — thanks @TheDemonTuan
- feat(api): add an opt-in
modelVisibilityAllowlist/modelVisibilityDenylistsettings pair to curate exactly which models/v1/modelsadvertises, mirrored into everyauto/*combo candidate pool so a denied model cannot be routed to via combo selection either (#11481 #11997) - feat(rankings): the Free Provider Rankings page shows what each provider actually served over the last 24 h. It ranked by ELO alone, which left a provider that answers every call with an error in first place; the usage data was already served by the API but never requested. A provider with too small a sample shows a dash, not a number (#11546 #11553) — thanks @maxmad64bis
- feat(guardrails): enforce a bounded, deterministic contract for Video Bridge transcripts — 256 cues, 4096 input code units and 4 KiB UTF-8 per cue, 64 KiB total text, malformed-Unicode rejection, focus-window scoping, cross-source reconciliation with contributing-source metadata, and a structural provenance trust boundary so caller JSON can never self-assert
embedded/audio-bridgeprovenance (#11652 #12009) - feat(video): orchestrate optional Video Bridge audio extraction and Audio Bridge STT behind a dual opt-in (operator setting AND per-request signal) — a new loopback-only broker
mode=audiooperation shares the frame path's exact process queue, deadline, AbortSignal, and byte budgets to extract a bounded mono 16 kHz PCM WAV from the same already-downloaded video, then reuses the existing Audio Bridge transcription boundary; provider segment timing is preserved when available and marked coarse otherwise, and every failure degrades to a visual-only-safe partial instead of throwing (#11654 #12012) - Add a tenant-bound Video Bridge drill-down lifecycle on top of the existing secure cache substrate: opaque hashed handles (never raw session/video identifiers), preview/standard/detail multiresolution variants resampled on read, response pagination capped at 8 frames and 32 MiB, and a new authenticated
/api/v1/video-bridge/drilldownconsumer route that stays disabled for remote access by default and denies cross-key access with the same response as a nonexistent handle (no existence oracle) (#12006) - test(video): Add the Video Bridge FU-07/FU-09 promotion-evidence harness — a frozen Zod manifest schema covering the 8 required scenario kinds (static scenes, rapid cuts, late facts, fades, blur, small text, close events, visual prompt injection) with a minimum of 3 repetitions per case, deterministic declarative fixture recipes (
videoBridgePromotionFixtures.ts), a pure medians/p95 metrics aggregator, a pure FU-07/FU-09 promotion-verdict evaluator applying the ticket's exact thresholds (missing token usage always holds), a digest-only persistence layer that never retains raw media or raw model… (#11656 #12008) - feat(video bridge): "embedded" transcript provenance can now be legitimately earned instead of merely asserted — a bounded, ...
v3.8.50
📊 Release by the numbers
| 👥 People who contributed | 248 |
| 📝 Commits in the cycle | 1,714 |
| 🔀 Pull requests referenced | 1,666 |
| 📋 Changelog entries | 1,182 |
| 🙌 Contributors credited in entries | 256 |
| 🤖 Automated dependency commits | 22 |
Entries by type
| Type | Count |
|---|---|
| 🐛 Fixes | 779 |
| ✨ Features | 169 |
| 📚 Docs | 29 |
| 🧹 Chore | 27 |
| 🧪 Tests | 15 |
| ♻️ Refactor | 5 |
| ⚡ Performance | 3 |
| providers | 2 |
| 🔒 Security | 2 |
| ⚙️ CI | 2 |
| deps | 2 |
| maint | 2 |
🏆 Top 25 contributors this cycle
By commits in ed2db6cb19..v3.8.50, author identities consolidated via .mailmap. Bots excluded.
| # | Contributor | Commits |
|---|---|---|
| 🥇 | diegosouzapw | 738 |
| 🥈 | backryun | 88 |
| 🥉 | Dizzle | 66 |
| 4 | Ravi Tharuma | 52 |
| 5 | Markus Hartung | 48 |
| 6 | Bob.Hou | 42 |
| 7 | Rouzbeh† | 38 |
| 8 | Xiangzhe | 31 |
| 9 | Paco Cartones | 28 |
| 10 | Nguyen Thanh Dat | 23 |
| 11 | Aman | 22 |
| 12 | Will Gordon | 19 |
| 13 | 小妍儿 ✨ | 17 |
| 14 | adevwithpurpose | 16 |
| 15 | Andrew B. | 10 |
| 16 | NOXX - Commiter | 10 |
| 17 | ignamiranda | 10 |
| 18 | Jonathan Bailey | 9 |
| 19 | Ke Jin | 9 |
| 20 | Austin Liu | 8 |
| 21 | Chewji | 8 |
| 22 | Prudhvi Vuda | 7 |
| 23 | benzntech | 7 |
| 24 | rinseaid | 7 |
| 25 | stanley | 7 |
Living section — regenerated 2026-08-12 from all cycle commits (cycle open ed2db6cb19 → tip). Bullets carry the merged PR and its author; direct pushes listed separately.
✨ New Features
-
feat(search): first-class X Search provider (
x-search) on… #10985 -
feat(core): add Layer A capability filter at router #5696
-
feat(providers): add DeepAI as paid API-key image provider #6671
-
feat(providers): add Naga.ac and… #6674 — thanks @chirag127
-
feat(api): add response content encoding verification —… #6736
-
feat(api): add plugins marketplace install endpoint with… #6752
-
feat(chatgpt-web): harden prompt-emulated… #7679 — thanks @horacecar
-
docs: add management authentication terminology guide #7786
-
feat(a2a): Conductor bridge — long-lived SSE consumer that… #8080
-
feat(a2a): the Agent Card (
/.well-known/agent.json) now… #8119 -
feat(dashboard): "Conductor" panel — OmniConductor fleet… #8221
-
feat(dashboard): Faro chat with voice on the Conductor panel —… #8222
-
feat(a2a): inbound delegation to the OmniConductor fleet — `POST… #8223
-
docs: add low-memory/small VPS optimization guide #8237
-
feat(providers): add connection-level… #8369 — thanks @Benson-mk
-
feat(copilot): add approval gate for runOmniRouteCli commands #8461
-
feat(ci): add windows-latest leg to test-bun-sqlite job #8468
-
feat(electron): Desktop app can now attach… #8799 — thanks @soulhakr
-
Database The
node:sqlitefallback now… #8870 — thanks @artickc -
Providers expands the Novita AI catalog… #8913 — thanks @jax-novita
-
feat(providers): native xAI Agent Tools passthrough on… #8964
-
feat(providers): add UnoRouter provider UnoRouter is an… #8978
-
feat(sse): deprecated the legacy
gemini-cli**upstream… #7034 #8980 -
Add a default-off connection setting for Codex, OpenAI, and…
-
Omit opaque encrypted reasoning values from persisted call logs… #9000
-
feat(providers): add Regolo AI OpenAI-compatible provider #9031
-
feat(db): add provider-scoped model aliases that survive… #9068
-
feat(cursor): surface a dismissible dashboard banner suggesting… #9173
-
feat(cursor): proactively renew Cursor sessions before their ~24h… #9173
-
feat(codex): accept parenthesized… #9208 — thanks @seakleangnhak
-
feat(usage): surface Claude thinking token… #9214 — thanks @luoyide
-
feat(ollama): add Ollama Local embedding… #9225 — thanks @HaoNgo232
-
feat(images): execute full combo strategy + fallback in… #9239
Adds open-sse/services/imageCombo.ts that expands combo targets, filters to images-capable, executes the priority strategy with handleImageGeneration per target, and returns the first success or last failure. Route patches detect combo names before model resolution and divert to the new execution path.
-
feat: make forwarded upstream response-header budget configurable… #9243
-
feat(providers): filter provider detail… #9247 — thanks @RobertsXML
-
feat(providers): make video_url… #9248 — thanks @HellFiveOsborn
-
feat(gemini): recursive type:object injection in schema… #9268
-
feat(dashboard): render a conditional "Get API key" link on the… #9270
-
feat(providers): accept JSON cookie… #9284 — thanks @AIB1TAL0S
-
feat(providers): support max reasoning effort for opencode-zen… #9318
-
feat(providers): expanded the NanoGPT (
nano-gpt.com) upstream… #9322 -
feat(sse): combo
system_messagesupports server-side… #5501 -
feat(sse): template expansion… #5501 #9414 — thanks @maxmad64bis
-
feat(sse): New-API/One-API/Sub2API aggregator balance detection… #9415
-
feat(catalog): added opt-in settings
hideAutoCombosand… #9418 -
feat(opencode-plugin): added… #9473 — thanks @omniroute
-
feat(providers): add native DeepSeek V4 Flash and Pro… #9485
-
feat(opencode-plugin): warm catalog startup from disk snapshot +… #9490
The config-shim hook now reads the last disk snapshot before fetching, so the provider registers immediately with the last-known-good catalog (~1-2s vs ~30s on a warm gateway). All six fetchers run concurrently via Promise.allSettled instead of sequentially. A failed refresh keeps the snapshot (no overwrite). An in-flight guard prevents concurrent refreshes for the same cache key. The features.diskCache: false opt-out disables the warm read entirely.
-
feat(models): Test All's "Auto-hide failed models" no longer… #9511
-
Add an advisory forgotten-sibling-tests report to pull-request… #9530
-
feat(providers): add Muse Code CLI provider preset #9544
-
feat(plugins): expose client request headers in plugin… #9570
-
feat(plugins): add onStreamComplete built-in event exposing… #9571
Adds a new
onStreamCompleteplugin event that fires after an SSE stream is fully
consumed, carrying usage token counts and timing metrics (latency, TTFT). Built-in
events now includeonStreamCompleteas a fire-and-forget lifecycle hook.Payload:
status,usage(prompt_tokens, completion_tokens, reasoning_tokens,
cache_read_input_tokens, cache_creation_input_tokens),timing(latencyMs, ttft),
model,provider,errorCode.Non-breaking — existing
onResponsehooks with{ streamed: true }remain unchanged. -
feat(audio): Soniox STT + TTS provider (
sx) — async… #9579 -
Show cache-read and cache-write token counts in request log rows and…
report them. (#9620) -
feat(memory): support custom OpenAI-compatible endpoints for… #9622
-
feat(resilience): add an opt-in watchdog for persistently slow… #9709
-
Onboarding: add an explicit, reviewable one-click setup for eligible…
with per-provider caution links, selectable confirmation, idempotent creation, and safe partial
retries. Existing provider connections are never changed and setup completion never enables
providers silently. (#9752) -
feat(settings): add a dedicated Modality Bridge settings page… #9782
-
feat(modality bridge): Transcribe chat audio for text-only… #9807
-
feat(memory):
PROVIDERS_SYSTEM_MUST_BE_FIRST(the… #6135 #7293 #9924 -
feat(api): API keys can disable prompt… #10001 — thanks @shixi-li
-
Add cliproxy provider exposure controls and manifest injection #7329 — thanks @KooshaPari
-
feat(infra): add a systemd autostart unit for Linux #8635
-
feat(db): add node sqlite adapter parity #8871 — thanks @epsilonode
-
feat(alibaba): free-tier routing with live quota sync #8893 — thanks @AndrianBalanescu
-
feat(oauth): add Raycast Pro provider with local auto-import #8895 — thanks @AndrianBalanescu
-
feat(executors): add isolated Claude Code bridge over Devin ACP #8914 — thanks @McLuck
-
feat: improve provider quota layouts #8916 — thanks @apoapostolov
-
feat(mcp): add omniroute_create_combo tool #8925 — thanks @lucasmellos
-
feat(ci): gate the publish on clean-install AND upgrade-over-previous #8953
-
feat(providers): add Conol (conol.ai) web session provider #8974 — thanks @artickc
-
Feat/combo provider wise model test #9011 — thanks @JoshimOfficial
-
feat(model-alias): add runtime Model Alias Resolver middleware #9020 — thanks @Egorich-print
-
feat(i18n): complete zh-CN localization for compression engines and dashboard UI #9038 — thanks @qianze0628
-
feat: Cheaper Inference provider (chat + native Responses + images, sponsor rail 2nd) #9043
-
feat(providers): add comprehensive support for self-hosted Firecrawl via FIRECRAWL_BASE_URL and custom base URLs #9052 — thanks @mad-gooze
-
feat(dahl): add manual API key option alongside auto-generated token #9077 — thanks @pizzav-xyz
-
feat(ci): G0 — reforça o trilho PR→release/ ** #9108
-
feat(providers): native xAI Agent Tools passthrough for /v1/responses #9111 — thanks @VXNCXNX
-
feat(.50): completa itens restantes — G13, G14, gap34, docs, R0.2 #9126
-
feat(g1): rewrite combo-strategy check to runtime-import approach #9131
-
feat(test:scoped): TIA-based local test runner (#8084 D1) #9143
-
feat(docker): publish next from active release branches #9181 — thanks @Zartharas
-
feat(usage): show Grok Build billing limits #9205 — thanks @xz-dev
-
feat(models): functional gateway mirrors + fix synced-substitution #9217
-
feat(admission): add adaptive overload protection for LLM routes #9262 — thanks @xz-dev
-
feat(i18n): update italian translations #9280 — thanks @Gecky2102
-
feat(dashboard): persist provider screen filters to URL for bookmarking #9307 — thanks @swingtempo
-
feat(api-manager): add provider-level model permissions #9313 — th...
Radar catalog export (rolling)
Export estável do catálogo OmniRoute para o Radar. Atualizado automaticamente; NÃO é um release de versão do produto.
v3.8.49
All 1383 entries from this cycle are listed below, one line each — descriptions are
trimmed to fit GitHub's 125,000-character release body. Full wording, context and links:
CHANGELOG.md.
Living section — regenerated 2026-07-19 from all 306 cycle commits (bump 2c62333 → tip). Bullets carry the merged PR and its author; direct pushes listed separately. Finalized at the v3.8.49 release.
✨ New Features
-
feat: generalize ensureThinkingBudget to all providers +… (#6979) — @rafaumeu
-
feat(6922): effort-tier aliases for glm-5.2 & mimo-v2.5 on… (#6987) — @rafaumeu
-
feat(providers): curated OpenRouter embeddings catalog + specialty merge… (#6994)
-
feat(quota): opt-in auto-ping to keep Codex quota windows warm (#6995)
-
feat(providers): add Agnes AI native provider support (#7035) — @HouMinXi
-
feat(sse): allow disabling
:comment heartbeats via… (#7036) — @xier2012 -
feat(perf): add performance.mark/measure to SSE pipeline +… (#7045) — @oyi77
-
feat(providers): add Dahl free inference provider (#7062) — @growab
-
feat(ci): boot-smoke the packed npm tarball (check:pack-boot,… (#7086)
-
feat(ci): hotfix fast-lane + tests-only E2E skip (WS3.1) (#7088)
-
feat(ci): continuous release-green — on-push quick gate + 3x/day… (#7089)
-
feat(ci): duration-balanced E2E shards via LPT bin-packing (WS4.1) (#7090)
-
feat(ci): TypeScript 7 native shadow for typecheck:core (WS4.2,… (#7091)
-
feat(release): npm staged publishing + pre-publish boot-smoke (WS1.3) (#7092)
-
feat(release): post-publish verifier — clean-container install + boot… (#7109)
-
feat(ci): Mergify merge queue + manual-train fallback runbook… (#7112)
-
feat(ci): Windows leg for Electron prepare smoke (WS1.5) (#7113)
-
feat(ci): Codecov patch coverage (informational) + fix missing… (#7114)
-
feat(sidecar): support conditional provider manifest refresh (#7130) — @KooshaPari
-
feat(homolog): real-environment E2E homologation suite (npm run… (#7133)
-
feat(usage): add Codex reset credit picker (#7154) — @JxnLexn
-
feat(ci): Trunk Flaky Tests uploads for vitest + Playwright E2E… (#7175)
-
feat(ci): Trunk Flaky Tests upload on the fast-path vitest job… (#7205)
-
feat(kiro): register GPT-5.6 Sol/Terra/Luna model family (#7209)
-
feat(dashboard): show Codex plan label in provider and quota views (#7210)
-
feat(dashboard): add reorder connections by availability button (#7211)
-
feat(dashboard): add 180D and 365D usage/cost analytics periods (#7213)
-
feat(api): add Vary: Accept-Encoding to token-authenticated /v1*… (#7217)
-
feat(api): expose GET /api/usage/model-latency-stats (#7218)
-
feat(dashboard): add compression-mode selector to Context & Cache combos… (#7219)
-
feat(sse): route GitHub Copilot Claude models through native… (#7223)
-
feat(mitm): add Antigravity reasoning-effort overrides (#7228)
-
feat: replace free-text model inputs with hidePaid-aware… (#7229)
-
feat: editable ComfyUI base-URL field + per-connection… (#7232)
-
feat(sse): add optional-enum null-omission idiom for codex… (#7233)
-
feat(sse): preserve tools/tool_choice for tool-bearing requests… (#7235)
-
feat(api): accept x-goog-api-key header for client-facing auth (#7236)
-
feat(sse): add native xAI Grok Imagine video generation provider (#7238)
-
feat: add Type filter and easiest-first sort to Free Provider… (#7240)
-
feat(cli): add Grok Build CLI tool setup (~/.grok/config.toml) (#7241)
-
feat(provider): add Chenzk API OpenAI-compatible gateway (#7246)
-
feat(providers): let custom connections opt into prompt-cache capability (#7257)
-
feat(db): include xp_audit_log in automatic retention/prune (#7260)
-
feat(api): structured X-Routing-Fallback-Reason header for relay… (#7262)
-
feat(compression): support RTK TOML schema v1 filters (#7281) — @JxnLexn
-
feat: add principal-scoped CCR MCP lifecycle (#7282) — @JxnLexn
-
feat(issue-agent): surface RecordedTriageTimeoutError as 504 (#7315) — @KooshaPari
-
feat(incident-response): structured incident response templates (#7334) — @KooshaPari
-
feat(providers): add xAI OAuth PKCE provider (#7399) — @fenix007
-
feat(models): advertise Claude reasoning-effort variants in /v1/models (#7497) — @thepigdestroyer
-
feat(kimi): sync Code, Web, and Moonshot providers (#7531) — @backryun
-
feat(resilience): guard OmniRoute peer routing loops (#7555) — @isiahw1
-
feat: add Mixedbread AI as embeddings provider (#7595)
-
feat(providers): add Rev AI speech-to-text provider (#7596)
-
feat: add Freepik (Magnific Mystic) image generation provider (#7597)
-
feat(sse): add DeepInfra as a video-generation provider (#7598)
-
feat(providers): add Felo chat-aggregator provider (#7599)
-
feat(sse): add Notion AI Web (Unofficial/Experimental) provider (#7600)
-
feat: add FreeTheAi as OpenAI-compatible gateway provider (#7602)
-
feat: add Gladia as an async speech-to-text provider (#7603)
-
feat: add EdgeTTS audio-tts provider (#7605)
-
feat(video): add Novita AI as video-generation provider (#7606)
-
feat: add Segmind image+video provider (#7608)
-
feat: add Microsoft Designer as image provider (#7609)
-
feat: per-model default reasoning_effort + no-think none on… (#7631)
-
feat(sse): per-model upstream header-response timeout override (#7632)
-
feat(dashboard): in-product guidance for prompt compression engines (#7634)
-
feat(usage): add TTFT/E2E-latency/tokens-per-second to model latency… (#7635)
-
feat: import providers from CSV/JSON file (#7636)
-
feat: confirm before removing a single connection (#7640)
-
feat(sse): honor excluded models in no-auth auto-combo candidate… (#7646)
-
feat(providers): add g4f.space no-key gateway… (#7647)
-
feat: rate-limit queue admission control (maxQueueDepth + 15s… (#7649)
-
feat(sse): generalize session affinity TTL to all providers (#7650)
-
feat: OpenRouter quota tracking (key/credits + free-window… (#7651)
-
feat(sse): quota tracking for AgentRouter, v0 (Vercel), FreeModel… (#7653)
-
feat(providers): Speechmatics STT, gTTS, VibeProxy preset (#6659, #6667,… (#7655)
-
feat(api): route Google AI Studio Imagen through… (#7656) — @danscMax
-
feat(auth): OIDC as optional dashboard admin login gate (password… (#6973) — @mikolaj92
-
feat(api): add pagination params to 8 DB modules + recharts… (#7046) — @oyi77
-
feat(proxy): operator-level proxy subscriptions (Karing-style) —… (#7299) — @xier2012
-
feat(grok-cli): align with official Grok Build client (#7358) — @backryun
-
feat(providers): Complete GHE Copilot OAuth provider implementation (#7546) — @hppsc1215
-
feat(guardrails): add CredentialMaskerGuardrail for API key/secret… (#7683) — @Securiteru
-
feat(perplexity): refresh provider integrations (#7687) — @backryun
-
feat(providers): notion-web live model discovery via getAvailableModels (#7696) — @artickc
-
feat(providers): add proactive cf_clearance/User-Agent hint to grok-web… (#7713)
-
feat: add live gRPC-web quota fetcher for grok-cli (#7714)
-
feat(api): add opt-in auto-sync scheduler for free-proxy sources (#7716)
-
feat(dashboard): show proxy name in badge, sort saved-proxy picker,… (#7720)
-
feat(cli): add auth export command for decrypted provider… (#7724)
-
feat(oauth): accept full ChatGPT session JSON for Codex manual import (#7725)
-
feat(sse): add nvidia NIM local RPM budget + concurrency cap (#7726)
-
feat(gemini-web): emulate OpenAI tool calling via the webTools prompt shim (#7727)
-
feat(services): introduce pluggable service-provider contract, migrate… (#7730)
-
feat(mitm): root-CA + per-host leaf certs for AgentBridge static… (#7731)
-
feat(providers): add hailuo-web (MiniMax web) chat provider (#7734)
-
feat: browser login for Grok Build provider (#7735)
-
feat(routing): wire interceptFetch tool interception into the chat… (#7736)
-
feat(sse): add X-OmniRoute-Decision routing trace header (#7765)
-
feat(providers): zai-web live model discovery with local-catalog fallback (#7766)
-
feat(api): sync upstream reasoning.supported_efforts into… (#7767)
-
feat(dashboard): pin Kimi providers first in category + official… (#7775)
-
feat(chaos+ponytail): parallel chaos-mode dispatch + ponytail output … (#7781) — @Moseyuh333
-
feat(perf): IC2 — cache provider connections by ID + lazy-decrypt… (#7787) — @oyi77
-
feat(quality): gate the free-tier headline so it can never silently… (#7798)
-
feat(providers): expose an explicit tier override for any provider… (#7838)
-
feat(routing): read-only auto/* candidate transparency + per-API-key… (#7839)
-
feat(catalog): map unmapped free tiers, add navy + aihorde, surface… (#7840)
-
feat(providers): add OpenRouter speech-to-text (audio transcription)… (#7861) — @Tasogarre
-
feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] (#7874) — @backryun
-
feat(providers): add 5 free-tier providers (ainative, aion, sealion,… (#7887)
-
feat(vnc-session): persistent noVNC browser login for web-cookie providers (#7892) — @Capslockb
-
feat(sse): add PromptQL playground provider (unofficial) (#7911) — @artickc
-
feat(cline): align ClinePass catalog and request protocol (#7914) — @backryun
-
feat: narrow mcp:connect scope + per-key HTTP tool-scope… (#7967)
-
feat: provider tab account search + mirrored top pagination (#7968)
-
feat: canonical numeric helpers + tier-1 (analytics) migration (#7969)
-
feat(sse): add HyperAgent (hyperagent.com) unofficial web provider (#7994) — @artickc
-
feat: copilot-m365...
v3.8.48
⚠️ Hotfix release. The published npm package for 3.8.47 crashed on every boot (#7065) and was deprecated — 3.8.48 is the first installable release of the v3.8.47 cycle, so everything listed under [3.8.47] below ships here.
🐛 Bug Fixes
- fix(build): ship
dist/head-response-guard.cjsin the npm tarball — the prepublish prune allowlist lacked it, so everyomnirouteboot of the published 3.8.47 crashed withERR_MODULE_NOT_FOUND(3rd occurrence of this class after tls-options/3.8.41); now allowlisted, enforced bycheck:pack-artifact, and guarded by a closure test that derives everyserver-ws.mjssibling import (#7065, #7040) - fix(build): Electron Windows packaging — the better-sqlite3 Electron-ABI rebuild now spawns
npx.cmdthrough a shell (Node's CVE-2024-27980 hardening made the shell-less spawn fail withstatus nullon Windows runners, breaking the v3.8.47 desktop build) - fix(ci): Sonar quality gate zeroed on new code — the coverage lcov now reaches the scanner at
coverage/lcov.info(it read 0% on every scan), the asyncisCloudEnabled()gate in the Kiro auto-import route is awaited (cloud sync ran even when disabled), the deadstructuredClonefallback in the reasoning-split clone is a real JSON fallback, the codex executor handles the asyncreader.cancel()rejection, deterministiclocaleComparesorts, a path-traversal guard inclassify-pr-changes.mjs, and the Docker better-sqlite3 rebuild uses npm's bundled node-gyp instead ofnpx --yes - chore(ci): the Sonar quality gate is informational (
sonar.qualitygate.wait=false) while the org's SonarCloud plan cannot associate the tuned "OmniRoute way" gate (coverage ≥60 aligned with the repo floor)
📦 Everything from the v3.8.47 cycle ships here
The 3.8.47 npm package was never installable (#7065), so 3.8.48 is the release that actually delivers the whole v3.8.47 cycle — full notes below:
- 9router Codex import: the Codex bulk-import endpoint (
POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt+ nestedproviderSpecificData), not just snake_case —normalizeCodexImportRecordmaps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard:tests/unit/codexBulkImport.test.ts(9router camelCase record, pre-suppliedproviderSpecificDatawithout an id_token, snake_case-not-overridden, and a full{accounts:[...]}flatten).
✨ New Features
-
feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)
-
feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)
-
feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)
-
feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)
-
feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)
-
feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)
-
feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)
-
feat(xai): route xAI clients to Grok's native
/v1/responsesendpoint. (#6709 — thanks @diegosouzapw) -
feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)
-
feat(release):
changelog.d/fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw) -
feat(quality):
validate-release-green --full-cireproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw) -
feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (
src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existingcommon.search/common.noResultsi18n keys (zero new locale edits) and the sharedInputicon="search"pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted intofilterSidebarSectionsByQuery()(src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard:tests/unit/sidebar-search-filter.test.ts,src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber) -
feat(combo):
auto/*combos gain a strict budget-cap fallback policy —X-OmniRoute-Budget-Fallback: strict(or the persistedconfig.budgetFallback: "strict") makes an over-budget request fail fast withHTTP 402instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existingX-OmniRoute-Budget/X-OmniRoute-Modeper-request controls (#6023/#6024/#6025), consolidated intoresolveRequestAutoControls(). Regression guard:tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470) -
Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)
-
Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the
#3587reasoning-modelmax_tokensbuffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev) -
feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (
comboStickyRoundRobinLimit, resolved viaresolveComboStickyRoundRobinLimit()— per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A newproviderStrategiessetting (Zod-validated map,src/shared/validation/settingsSchemas.ts) lets a specific provider override the globalfallbackStrategy/stickyRoundRobinLimitwithout touching the account-wide default, wired intogetProviderCredentials()(src/sse/services/auth.ts) ahead of the global fallback. Regression guard:tests/unit/combo-rr-sticky-9router.test.ts,tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen) -
feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG →
@lobehub/iconsReact components →thesvg.orgCDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard:tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7) -
Skill Collector CLI detection: new
GET /api/skills/collect/detect+POST /api/skills/collect/install(and thecli-skill-collectoragent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally viagetCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes andGET/POST /api/github-skillsnow require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES+SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). Theomniroute_github_skills_installMCP tool now reports the honestaction: "planned"instead of"installed", matching the REST route (#6294 — thanks @Moseyuh333) -
ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to
cp(matching theOAUTH_PROVIDERScatalog alias) so<alias>/<modelId>routing resolves correctly, the OAuth refresh dispatch now routesclinepassto the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard:tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok) -
feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no
aorAAAAAG-prefixed refresh token) a...
v3.8.47
- 9router Codex import: the Codex bulk-import endpoint (
POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt+ nestedproviderSpecificData), not just snake_case —normalizeCodexImportRecordmaps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard:tests/unit/codexBulkImport.test.ts(9router camelCase record, pre-suppliedproviderSpecificDatawithout an id_token, snake_case-not-overridden, and a full{accounts:[...]}flatten).
✨ New Features
-
feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)
-
feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)
-
feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)
-
feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)
-
feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)
-
feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)
-
feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)
-
feat(xai): route xAI clients to Grok's native
/v1/responsesendpoint. (#6709 — thanks @diegosouzapw) -
feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)
-
feat(release):
changelog.d/fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw) -
feat(quality):
validate-release-green --full-cireproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw) -
feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (
src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existingcommon.search/common.noResultsi18n keys (zero new locale edits) and the sharedInputicon="search"pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted intofilterSidebarSectionsByQuery()(src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard:tests/unit/sidebar-search-filter.test.ts,src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber) -
feat(combo):
auto/*combos gain a strict budget-cap fallback policy —X-OmniRoute-Budget-Fallback: strict(or the persistedconfig.budgetFallback: "strict") makes an over-budget request fail fast withHTTP 402instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existingX-OmniRoute-Budget/X-OmniRoute-Modeper-request controls (#6023/#6024/#6025), consolidated intoresolveRequestAutoControls(). Regression guard:tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470) -
Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)
-
Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the
#3587reasoning-modelmax_tokensbuffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev) -
feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (
comboStickyRoundRobinLimit, resolved viaresolveComboStickyRoundRobinLimit()— per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A newproviderStrategiessetting (Zod-validated map,src/shared/validation/settingsSchemas.ts) lets a specific provider override the globalfallbackStrategy/stickyRoundRobinLimitwithout touching the account-wide default, wired intogetProviderCredentials()(src/sse/services/auth.ts) ahead of the global fallback. Regression guard:tests/unit/combo-rr-sticky-9router.test.ts,tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen) -
feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG →
@lobehub/iconsReact components →thesvg.orgCDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard:tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7) -
Skill Collector CLI detection: new
GET /api/skills/collect/detect+POST /api/skills/collect/install(and thecli-skill-collectoragent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally viagetCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes andGET/POST /api/github-skillsnow require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES+SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). Theomniroute_github_skills_installMCP tool now reports the honestaction: "planned"instead of"installed", matching the REST route (#6294 — thanks @Moseyuh333) -
ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to
cp(matching theOAUTH_PROVIDERScatalog alias) so<alias>/<modelId>routing resolves correctly, the OAuth refresh dispatch now routesclinepassto the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard:tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok) -
feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no
aorAAAAAG-prefixed refresh token) and can't refresh through the AWS OIDC/Kiro-social path, sotryAwsSsoCache()now detects them (authMethod/provider === "externalidp") and refreshes via the org IdP's owntokenEndpoint(public-client OAuth2 refresh grant, no client secret), persistingTokenType: EXTERNAL_IDPgating so the runtime executor sends the header the AWS CodeWhisperer API requires for these accounts;tokenEndpointis SSRF-guarded against an HTTPS + known-IdP-host-suffix allowlist. (#6363 — thanks @artickc) -
Kiro long-lived API key auth: new
/api/oauth/kiro/api-keyroute +KiroService.validateApiKeylet a Kiro account be linked with a long-lived AWS CodeWhisperer/Kiro API key instead of the interactive OAuth device flow, with live per-account model discovery (ListAvailableModels, 5-minute cache) layered over the existing static registry fallback (#6587 — thanks @strangersp) -
Chaos Mode: multi-model parallel/collaborative task execution — dispatches a task to every active provider connection at once (parallel) or chains outputs sequentially so each model builds on the previous one's answer (collaborative), configurable via Dashboard → Chaos Mode (
GET/PUT/DELETE /api/chaos/config) and gated per-API-key via a newchaosModeEnabledpermission (opt-in — disabled by default globally and per key).POST /api/chaos/run(dashboard session) andPOST /api/skills/collect/chaos(external Bearer-token) delegate to a sharedexecuteChaosRun()engine (src/lib/chaos/chaosExecutor.ts) that dispatches in-process via the established synthetic-Request/route-handler pattern (no network hop, no hardcoded port), with a concurrency cap (max 10 parallel), configurablemax_tokens(256–128k), a clear error whenstreamis requested, and collaborative-chain info (provider order + input size). Fixes external Bearer-auth bypass and stale config-cache leakage. Regression guard:tests/unit/chaos-config.test.ts,tests/unit/chaos-executor.test.ts,tests/unit/chaos-api-routes.test.ts. (#6728 — thanks @Moseyuh333) -
feat(cli): 2 new...
v3.8.46
✨ New Features
- feat(sse): hide paid-only models from
auto/*routing whenhidePaidModelsis on (#6512) — follow-up to #6328/#6495. PR #6495 hid paid-only models from theGET /v1/modelslisting, butauto/*combos (auto/best-coding,auto/glm, …) could still pick a paid-only backend into their candidate pool → a 402/403 at request time.createVirtualAutoCombonow filters the candidate pool through the new pureopen-sse/services/autoCombo/paidModelFilter.ts(filterPaidOnlyCandidates), applying the same free-model predicate #6495 uses incatalog.ts(providerHasFreeModels(provider) && isFreeModel(provider, {id})) wheneversettings.hidePaidModels === true. Applied before the category/tier/family narrowing, so it covers everyauto/*combo; an all-paid pool degrades to the existing graceful empty-pool path. Opt-in — default OFF leaves the pool unchanged (identity). Regression guard:tests/unit/autoCombo/paid-model-filter-6512.test.ts(4, incl. the default-off identity guard). - feat(sse): provider-family auto combos —
auto/glm,auto/minimax,auto/mimo,auto/zai,auto/gemma,auto/llama,auto/gemini(#6453) — new routable ids that materialize an on-demand virtual combo spanning whatever installed backends currently expose that model family, degrading gracefully as backends rotate. A new pureopen-sse/services/autoCombo/modelFamily.ts(detectModelFamily) classifies by model-id prefix for six families;zaiis instead resolved by provider id (z.ai's hosted API serves the sameglm-*model ids as every other GLM backend, soauto/zaimeans "route to my z.ai backend specifically" vsauto/glm's "any connected GLM backend"). Reuses the existingcreateVirtualAutoComboon-demand materialization path (no DB writes) and the/v1/modelscatalog advertising loop. Regression guard:tests/unit/autoCombo/provider-family-combos.test.ts(11). - feat(proxy): native proxy-pool round-robin / egress IP rotation (#6365) — a scope (global / provider / account) can now hold multiple proxies as a pool with a rotation strategy, so outbound requests cycle their egress IP instead of pinning one proxy per scope. Migration
117_proxy_pool_rotation.sqllifts theUNIQUE(scope, scope_id)constraint (rebuild via the canonical rename/copy/drop; existing single assignments become 1-element pools) and adds aproxy_scope_rotationcompanion table holding the per-scope strategy + a persisted monotonic round-robin cursor. Strategies:round-robin(default, monotonic cursor — neverMath.random),random, andsticky-per-N-min. Resolution (resolveProxyForScopeFromRegistry/resolveProxyForConnectionFromRegistry) now fetches the alive, position-ordered candidate set (unchangedPROXY_ALIVE_PREDICATE) and applies the strategy; an empty / all-dead pool still returnsnull— the #6246 fail-closed guard is untouched (never falls through to direct egress). Backend + DB only; dashboard pool-builder UI is a follow-up. Regression guard:tests/unit/proxy-pool-rotation-6365.test.ts(8, incl. fail-closed + backward-compat). - feat(providers): end-to-end tool/function calling on the native Gemini
/v1betaendpoint (#6222) — both directions of the Gemini↔OpenAI conversion now preserve tool data (previously silently dropped). Request side:convertGeminiToInternal(extracted to its own testable module) mapstools[].functionDeclarations→ OpenAItools, priorfunctionCallparts → assistanttool_calls, andfunctionResponseparts →tool-role messages. Response side:convertOpenAIResponseToGeminiemitsparts[].functionCall {name,args}frommessage.tool_calls, and the streamingopenAIChunkToGeminiChunkaccumulates fragmentedtool_callsdeltas by index into completefunctionCallparts. The non-Gemini client paths (Claude, OpenAI-Responses) already preserved tool calls — this closes the gap specific to the native Gemini surface. Regression guard:tests/unit/v1beta-gemini-tool-calling-6222.test.ts(6, incl. a streaming SSE round-trip). - feat(providers): copilot-m365-web enterprise / work tier support (#6334) — mirrors the EDU-tier pattern (#6210):
M365ConnectionParamsgains anagentfield, a new opt-inM365_ENTERPRISE_OVERRIDESpreset (agent=work,scenario=officeweb,licenseType=Premium) applies viaproviderSpecificData.tier="enterprise"(alias"work"), andagentis also overridable directly viaproviderSpecificData.agent.buildWsUrlwas hardcodingagent="web"(the one enterprise-distinguishing param with no override path), so a Premium work account handshook then returned an empty stream. The individual and EDU paths are untouched. Kilo's dup flag vs #6210 (EDU tier) was a false positive — different tier. Regression guard:tests/unit/copilot-m365-enterprise-6334.test.ts(7). End-to-end confirmation on a real Premium work account is a live-VPS validation follow-up (Hard Rule #18). (thanks @Forcerecon) - feat(api): standardized, provider-agnostic
effort+thinkingrequest params (#6241) — a thin standardization layer over the existing mature per-provider reasoning plumbing (no provider mapper touched).providerChatCompletionSchemagains a canonicaleffort(reusing the sharednone/low/medium/high/xhighvocabulary — the UI tiersextra/maxcollapse ontoxhigh) and a booleanthinking. A purenormalizeReasoningRequest(wired once insrc/sse/handlers/chat.ts, before any reasoning field is read) folds them onto the fields the translators already consume (reasoning_effort/reasoning.effort/thinking), so they fan out to Anthropic / Gemini / xAI / Responses — an explicit clientreasoning_effort/ object-shapedthinkingalways wins (backward-compatible)./modelsadditively exposessupportsThinking+effort_tiersso the frontend can render the toggles (UI component is a follow-up). Regression guard:tests/unit/effort-thinking-standardization-6241.test.ts(12). (thanks @Iammilansoni, @shabeer) - feat(combo): new
pipeline(sequential) combo strategy (#6297) — the 18th routing strategy runs targets in order, threading each step's output into the next step's input, with an optional per-stepprompt(system instruction); only the final step's response is returned. Distinct fromfusion(parallel fan-out + judge). Implemented as a self-containedopen-sse/services/pipeline.ts(sibling tofusion.ts), dispatched fromcombo.ts; the step list reusescombo.modelsorder and reads an optionalpromptoff each target (backward-compatible — ignored by every other strategy). Intermediate steps run non-streaming with tools stripped (complete prose to thread forward); the final step keeps the client'sstreamflag + tools. A failing/empty/unparseable intermediate step fails the whole pipeline explicitly via a sanitized error (never silently swallowed). Kilo's dup flag vs #563 was a false positive (that's model→chain selection; this is a sequential chain). Regression guard:tests/unit/combo-pipeline-strategy.test.ts(5). (thanks @ofekbetzalel) - feat(ci):
check:test-maskingnow flags inline-reimplemented prod conditions (#6348) — a new report-only subcheck (v2, 6A.10 family) catches the wrong-shape contract test: a test that recomputes the condition under test inline instead of importing/exercising the real function (the #6216 class, where=== 500→>= 500stayed green because the test re-implemented the branch). For each added/modified test file it warns when the file textually duplicates a ≥3-token conditional from a production file touched in the same PR and does not import the symbol/module owning it, via a pure, fixture-testedfindReimplementedConditions()with an allowlist mirroringassertReductionAllowlist. Report-only for now (does not fail the gate) — to be promoted to blocking after a triage cycle. Regression guard:tests/unit/check-test-masking.test.ts(45). - feat(sse): per-connection routing override (native vs CLIProxyAPI) (#6339) — the previously-dead
isCliproxyapiDeepModeEnabledhelper is now wired intoresolveExecutorWithProxy: a single connection can opt itself into the CLIProxyAPI passthrough executor viaproviderSpecificData.cliproxyapiMode="claude-native", with precedence connection override > providerupstream_proxy_configmode > default.resolveExecutorWithProxynow receives the resolved connection'sproviderSpecificData(threaded fromchatCore.ts), so one connection can deep-route while the provider's other connections stay native — no DB schema change (the toggle rides inproviderSpecificData). Also resolves the same-provider mixing ask in #6340. Regression guard:tests/unit/chatcore-executor-proxy.test.ts(9). (thanks @RaviTharuma) - feat(dashboard): "Add session cookie" modal now shows a prominent "Open ‹host› →" link to the provider's own site (#6268) — every
-webcookie-session provider (chatgpt-web, claude-web, gemini-web, kimi-web, lmarena, qwen-web, m365-copilot-web, …) renders a one-click external link (opening the provider's login/home page in a new tab) so operators no longer tab away to retype the URL mid-setup. The host resolves from a pure, unit-testedresolveWebProviderHost()(prefersWEB_COOKIE_PROVIDERS[id].website, falls back to the registrybaseUrlorigin); non-web providers render e...
v3.8.45
✨ New Features
- feat(providers): add Yuanbao (web) as a cookie-session provider (#6196) —
yuanbao-web(Tencent Yuanbao,yuanbao.tencent.com) with cookie-only auth (hy_user/hy_token+ public agent id), SSE→OpenAI translation incl.reasoning_content, exposing DeepSeek V3/R1 + Hunyuan / Hunyuan-T1. Regression guard:tests/unit/providers-yuanbao-web.test.ts.together-webwas deferred (no verifiable web-session endpoint — needs a captured request) andhuggingchat-webdropped (the existinghuggingchatalready is a web-cookie provider). (thanks @chirag127) - feat(providers): route the built-in agentrouter through the dynamic Claude-Code wire image (#6056) — a small static allow-set (
CC_WIRE_IMAGE_BUILTINSinopen-sse/services/ccWireImageBuiltins.ts), consulted byisClaudeCodeCompatible/isClaudeCodeCompatibleProvider/applyFingerprint, makes agentrouter adopt the CC wire-image headers + fingerprint while guarding the CC baseUrl/auth branches so it keeps its own registrybaseUrlandx-api-keyauth. Regression guard:tests/unit/agentrouter-cc-wire-image.test.ts(asserts the wire image is applied AND agentrouter's baseUrl/auth are preserved). Live WAF-acceptance against agentrouter.org is a VPS validation follow-up (Hard Rule #18). - feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174) —
cloudflare-aiis removed from the bulk-add exclusion list and the bulk parser gains a 3-fieldname|accountId|apiKeymode; the bulk route now builds a per-entryproviderSpecificDataso each key carries its ownaccountId(fixing the previous shared-object reuse), and both the create + key-validation paths receive it. Regression guard:tests/unit/bulk-api-key-parser-cloudflare.test.ts. (thanks @muflifadla38) - feat(dashboard): routing/settings UX clarity (#6147) — (1) weighted combos show the effective routing share % next to each weight when weights don't sum to 100 (
WeightTotalBar.tsx); (2) the status widget's user-facing "Cloud Sync" label is renamed to "Remote Settings Sync" (CloudSyncStatus.tsx; internal ids/state untouched); (3) built-in providers gain an opt-in advanced base-URL override (isBaseUrlOverrideEligibleProvider, hidden behind an "Advanced" toggle, reusing the existingproviderSpecificData.baseUrlpersistence — not globally widened). Regression guard:tests/unit/routing-settings-ux-6147.test.ts. - feat(combo): add an option to disable session stickiness, per-combo or globally — round-robin / random combos can rotate to a different connection on every request instead of pinning a whole conversation to one connection by its first-message hash. Resolution precedence per-combo
config.disableSessionStickiness→ globalsettings.disableSessionStickiness→ defaultfalse(preserves the #3825 prompt-cache/504 fix); gates both stickiness call sites inopen-sse/services/combo.ts. Exposed as a global toggle (Combo Defaults) and a per-combo Inherit/on/off control. (#6168) Regression guard:tests/unit/combo-disable-session-stickiness.test.ts. (thanks @RCrushMe) - feat(docker): add the
OMNIROUTE_NO_SUDOenv flag for root-less / user-namespaced deployments — the MITM cert-trust command path (resolveSudoSpawninsrc/mitm/systemCommands.ts) now strips the leadingsudowhen the flag is truthy, in addition to the existing root / sudo-missing cases, so the Proxy Agent runs withoutsudo(the operator trusts the CA manually, e.g. viaNODE_EXTRA_CA_CERTS). Argv-arrayspawnpreserved — no shell interpolation (Hard Rule #13). (#6122) Regression guard:tests/unit/mitm-systemCommands-no-sudo.test.ts. (thanks @powellnorma) - feat(providers): add Requesty as an OpenAI-compatible gateway provider (BYOK, base
https://router.requesty.ai/v1, ~200 free requests/day) — wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/requesty/,src/shared/constants/providers/apikey/gateways.ts). (#6120) Regression guard:tests/unit/requesty-provider.test.ts. (thanks @chirag127) - feat(dashboard): add configured-only / available-only filters to the Free Provider Rankings page (#6150) — hide providers you haven't configured, or whose connections are all rate-limited / out of quota, via server-side query params (
?configuredOnly/?availableOnlyonGET /api/free-provider-rankings) backed by a testable lib helper reusing the in-process connection state (no Redis). Both filters default off, so the default view is unchanged; this supersedes the earlier client-side "Configured Only" toggle (#6245) with an available-only dimension and unit-tested logic. Regression guard:tests/unit/freeProviderRankings-filters.test.ts. - feat(rankings): add a 'Configured Only' filter to the Free Provider Rankings page, so the table can be narrowed to just the providers you have configured connections for (with an empty-state hint when none are configured). New
en.jsonkeys and a pure filter helper covered bytests/unit/free-provider-rankings-configured-filter.test.ts. (#6245, closes #6150 — thanks @Iammilansoni)
🔧 Bug Fixes
- fix(mitm): the test suite and CI can never mutate the OS trust store again —
OMNIROUTE_SKIP_SYSTEM_TRUST=1(set by the global test setup and all CI workflows) makesinstallCert/uninstallCert/installTproxyCaskip the privileged OS dispatch while preserving the #4546 environment-skip contract. Root cause of the self-hosted runner incident: a cert-flow integration test installed a 105-byte fake PEM into/usr/local/share/ca-certificates, breaking ALL system TLS on the VM. Regression guard:tests/unit/system-trust-test-guard.test.ts. (#6310) - fix(security):
/api/keys/{id}/devicesanswers a clean method-first 405 for undocumented HTTP methods (e.g. the newQUERY) via a dedicatedhttp-method-guardrule — the auth layer was answering 401 first, failing schemathesis's unsupported-methods check. Same pattern as the v3.8.44 TRACE fix. Regression guard:tests/unit/dast-method-not-allowed.test.ts. - fix(combo): the #6216 empty-stream failover is restricted to truly empty bodies (zero bytes — the Gemini HTTP-200-empty case), restoring the #3399/#3685 pass-through contracts for
[DONE]-terminated empty streams and incomplete Claude lifecycles. New guard:#5976 truly EMPTY streaming body → invalid for combo failover(87/87 across both suites). - fix(combo): 5 streaming-path fixes — locked-stream 500, error-frame-only-if-no-content, Gemini
MALFORMED_RESPONSE→content_filter failover, correlationId substring search, per-model-500 lockout skip + request-logger UI detail. Maintainer follow-up:releaseQualityClonecancels the abandoned quality-check tee branch (per-request memory) + regression test. (#6216 — thanks @hartmark) - fix(skills): generate the missing
omni-github-skillsregistry entry (the #6186 catalog addition never ran the generator — 8 integration assertions split between old/new counts) and align the agent-skills catalog counts across integration + unit suites (43 = 23 API + 20 CLI; 44 with config). - fix(a2a): finish the #6186 catalog-count update —
listCapabilitiesmetadata reportedcoverage.api.total: 22(type literal + value) andSkillCoverageSchemapinnedz.literal(22), so the schema would REJECT the correct runtime value with 23 API skills. All three aligned to 23. - fix(github-skills): add a missing import, unit tests and a settings JSON-parse fix for the GitHub agent-skill discovery/import flow. (#6186 — thanks @Moseyuh333)
- fix(api):
POST /api/github-skillsvalidates its body with a Zod schema (validateBody) instead of blindrequest.json()destructuring — a non-arraytargetswould crash.map. Regression guard:tests/unit/github-skills-route-validation.test.ts. - fix(docker): add
id=to the BuildKit cache mounts so strict builders (e.g. buildkitd with strict frontend parsing) accept the Dockerfile. (#6291 — thanks @karimalsalah) - fix(oauth): register
zedin the OAuthPROVIDERSmap (fixes "Unknown provider" on the Zed sign-in flow) (#6078 — thanks @anki1kr), and alignzedinOAUTH_PROVIDER_IDS+ the config enum after the merge. - fix(doubao-web): switch the Doubao web provider to the Dola global endpoint. (#6235 — thanks @backryun)
- fix(doctor): resolve two false-positive WARNs in the doctor diagnostics (#6163, closes #6162 — thanks @arssnndr)
- fix(providers): refresh the GitHub Copilot model catalog to the current upstream set. (#6154 — thanks @backryun)
- fix(providers): correct the Kiro model catalog to real upstream ids — fabricated
claude-opus-4.7/claude-sonnet-4.6entries removed, realclaude-sonnet-5/claude-sonnet-4.5/claude-haiku-4.5kept. (#6170...
v3.8.44
✨ New Features
- feat(resilience): throttle upstream quota fetches on the per-request preflight path (#6009) — a new global min-interval gate (
open-sse/services/quotaFetchThrottle.ts) spaces the actual network calls made by the Codex quota fetcher so that many accounts on one IP no longer fetch quota in the same second (which, perrouter-for-me/CLIProxyAPI#2385, can get a Codex OAuth token revoked). Complements the existing bulk-sync spacing (PROVIDER_LIMITS_SYNC_SPACING_MS) which already serialized the periodic provider-limits sync — this covers the concurrent combo/preflight path it didn't. Cache hits are never delayed; fail-open (only ever awaits a timer). Configurable viaOMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS(default 250ms, clamped 0..5000;0disables). Regression guard:tests/unit/quota-fetch-throttle-6009.test.ts(5). (thanks @powellnorma) - feat(autoCombo): add per-request Auto-Combo controls via two headers (#6024 / #6025 / #6023) —
X-OmniRoute-Modesteers anautocombo's scoring for a single request (friendly presetsfast/balanced/quality/cheap/reliable/offlineor a raw mode-pack name;balancedforces the default weights), andX-OmniRoute-Budgetsets a hard per-request USD cost ceiling. Both override the combo's stored config only for the request that carries them; unknown/garbage values are ignored so the saved config is preserved. The resolvers are pure (open-sse/services/autoCombo/requestControls.ts) and feed the engine's existingconfig.modePack/config.budgetCapinputs — no engine changes. Regression guard:tests/unit/auto-combo-request-controls-6024.test.ts(5). (thanks @chirag127) - feat(providers): add the Kenari OpenAI-compatible gateway (BYOK). Regression guard:
tests/unit/kenari.test.ts. (thanks @doedja) - feat(models): add
claude-sonnet-5to the Antigravity model catalog (alias mapping inantigravityModelAliases.ts) (#6103). Regression guard:tests/unit/antigravity-model-aliases.test.ts. (thanks @anki1kr) - feat(api): add
/v1/ocrendpoint (Mistral OCR), an OCR provider category, and Mistral moderation support. (#5950) (thanks @waguriagentic) - Discovery tool (Phase 2): add the
discoveryResultsDB module (CRUD over thediscovery_resultstable, migration 074) and wire the opt-in provider-discovery service to persist and read findings through it (persistDiscoveryResult,getDiscoveryResults,getDiscoveryResultById,markVerified,deleteDiscoveryResult) with(provider, method, endpoint)upsert de-duplication. Adds the/api/discovery/*HTTP surface —GET /results,GET|DELETE /results/:id,POST /scan,POST /verify/:id— under strict loopback-only authorization (/api/discovery/is inLOCAL_ONLY_API_PREFIXESand is NOT manage-scope-bypassable, so thescanroute's outbound probes can never be reached from a tunnel/remote origin). Adds a dashboard UI tab (Tools → Discovery,/dashboard/discovery) to run scans and review, verify, or delete findings. The service stays opt-in / default-off. (#5939) - feat(api): expose a read-only provider plugin manifest at
GET /api/v1/provider-plugin-manifestfor sidecar/relay discovery. (#6001) (thanks @KooshaPari) - feat(sidecar): advertise the provider manifest URL to Bifrost/CLIProxyAPI via the
X-OmniRoute-Provider-Manifest-Urlheader (OMNIROUTE_PROVIDER_MANIFEST_URL). (#6007) (thanks @KooshaPari) - feat(autoCombo): add a latency/speed-optimized routing mode (shared
rankBySpeedscoring core) plus theomniroute_pick_fastest_modelMCP tool. (#6011) (thanks @KooshaPari) - feat(resilience): surface Codex banked reset credits per connected account (#5199) — the Codex quota parsers (
buildCodexUsageQuotas,parseCodexUsageResponse) now additively readrate_limit_reset_credits.available_count(+ optionalrate_limit_reached_type) from the/wham/usagepayload OmniRoute already fetches, and the provider-limits dashboard renders a "Banked Reset Credits" row when a positive count is present. Display-only and fail-open — the field is eligibility-gated, so accounts without it are unaffected (parsers never throw on absent/garbage shapes); redemption (an unofficial mutating endpoint) is intentionally out of scope. Regression guard:tests/unit/codex-banked-reset-credits-5199.test.ts(8). (thanks @ofekbetzalel) - feat(providers): add sign-up geo-restriction notices for SenseNova and StepFun (#5462) — the provider add-form now warns that SenseNova's console appears to require a Chinese (+86) phone number with no documented international path, and that StepFun's default endpoint is its China platform while a global StepFun Open Platform (
platform.stepfun.ai, operated by Sparkling AI Pte. Ltd., Singapore) with email/Google/Discord login exists for international users. Informationalnoticeonly — neither provider is disabled. Regression guard:tests/unit/regional-provider-cn-notices-5462.test.ts. (thanks @chirag127) - feat(usage): add on-demand period-scoped usage-data reset (Settings → System Storage) with a purge API and time-window selector. (#5831)
- feat(claude-code): add an opt-in auto-permission classifier compat mode (off/auto/always) for Claude Code, toggleable from the CLI Code settings. (#5810)
- feat(providers): add optional client-identity header profiles for compatible nodes — preset User-Agent/fingerprint headers (e.g. matching a known CLI) merged into the existing customHeaders field. (#5812)
- feat(build): add a backend-only fast build mode (
scripts/build/build-next-isolated.mjs+backendOnlyPages.mjs) that skips compiling the dashboard frontend pages, cutting local/CI build time for backend-only changes. (#6119 — thanks @artickc) - feat(minimax): extract MiniMax M3's raw
<think>...</think>leakage intoreasoning_contenton the 8 OpenAI-format provider tiers, leaving the Claude-formatminimax/minimax-cntiers untouched (they already report reasoning correctly). (#6073 — thanks @KooshaPari) - feat(services): promote Bifrost (
@maximhq/bifrost— Go AI-gateway) from an env-only relay sidecar to a first-class embedded/supervised service, matching the existing cliproxy/9router model — installer, bootstrapSERVICES[]entry, migration 113 DB seed, 7 lifecycle API routes under/api/services/bifrost/(loopback-only), a dashboard tab, and relay auto-wiring that defaultsBIFROST_BASE_URLto the supervised port when running. Implements item #2 of #5670; the broader RouterBackend contract (items #1, #3-#5) stays out of scope. (#5817, part of #5670) - feat(services): add Mux (
coder/mux— local agent-orchestration daemon) as a fourth-tier embedded service on the existingServiceSupervisorframework — npm-based installer,bootstrap.tsregistration, migration 113 DB seed, 7 lifecycle API routes under/api/services/mux/(loopback-only, defense-in-depth bind to 127.0.0.1), and a dashboard tab reusing the shared service-management components. (#6034) - feat(xai): surface Grok/xAI usage on the quota dashboard via local
usageHistoryaggregation (getXaiUsage) — since xAI exposes no per-account quota API, this sums tokens routed to the connection fromusage_historyand reports them as a cumulative, uncapped quota, mirroring the existing Xiaomi MiMo self-track pattern. (#5806) - feat(minimax): extract MiniMax M3's raw
<think>...</think>tags into a separatereasoning_contentfield on the 8 provider tiers that register M3 withformat:"openai"(trae, huggingchat, bazaarlink, ollama-cloud, opencode, cline, opencode-zen, codebuddy-cn) — previously the thinking text leaked directly intocontent. Reuses the existingextractThinkingFromContentprimitive, extending its allowlist with a minimax-m3-only pattern; the two direct minimax/minimax-cn tiers are untouched since they already surface reasoning natively over Anthropic's Messages format. (Inspired by 9router#2231.) (#6050 — thanks @KooshaPari) - feat(i18n): auto-detect the browser language on first visit — a pure
detectBrowserLocale()matcher (exact match,zh-HK/zh-MOfolded tozh-TW, language-prefix match, elsenull) plus a client-onlyLocaleAutoDetectcomponent mounted once in the root layout. When no locale cookie is set yet, it readsnavigator.languages, computes a match against the supported locales, and persists it via the same cookie/localStorage writerLanguageSelectoralready used (extracted toshared/lib/persistLocale.ts). (Inspired by 9router#1324.) (#5979) - feat(cli-tools): add CodeWhale — the actively-maintained successor to DeepSeek TUI (same author, renamed pr...
v3.8.43
[3.8.43] — 2026-07-02
✨ New Features
-
usage (quota percentages + provider USD drilldown):
@@om-usageand the HTTP usage endpoint now report personal API-key quotas as remaining percentages (USD amounts stay out of the command output), provider quota remaining is scaled by the configured quota cutoff so the protected reserve reads as 0% left, and the quota dashboard regains a provider USD cost drilldown (/api/usage/provider-window-costs+ProviderUsdCostModal, management-auth gated). Also honors observed provider quota resets: a same-resetAtreset (usage dropping back to the reset floor) is detected and preferred over stale recorded weekly events for provider USD windows and API-key USD quotas. Newsrc/lib/usage/providerWindowCosts.ts. Regression guards:tests/unit/provider-window-costs.test.ts,tests/unit/internal-usage-command.test.ts,tests/unit/api-key-usage-limits.test.ts,tests/unit/lib/quota-reset-events.test.ts. Extracted from #5863 by @Witroch4. -
dashboard (live WS behind reverse proxy): the live dashboard WebSocket can now be fronted by a reverse proxy or Cloudflare Tunnel via
NEXT_PUBLIC_LIVE_WS_PUBLIC_URL(e.g.wss://ws.my-ai.com/live-ws). The URL is honored both at build time (env inlined into the bundle) and at runtime for prebuilt Docker/npm images: the/api/v1/ws?handshake=1handshake now echoes a lazily-readlive.publicUrl(onlyws:///wss://values are accepted; anything else is rejected tonull), anduseLiveDashboardresolves the URL from that handshake before connecting, falling back to the previousws(s)://hostname:20129default. Also documentsLIVE_WS_ALLOWED_HOSTSand aligns the GitLab Duo OAuth scopes line in.env.examplewith the live config (ai_features read_user). Regression guard:tests/unit/live-ws-public-url.test.ts(5). (#5877 by @ianriizky) -
providers (CLI profile auto-sync): opt-in toggles to auto-regenerate CLI tool profiles after a provider model sync. When enabled, a model-catalog change (re)writes that tool's profile files from the live catalog — Codex (
~/.codex/*.config.toml) and now Claude Code (~/.claude/profiles/<name>/settings.json, via an extractedsyncClaudeProfilesFromModels+ a newclaudeProfileAutoSync.tsmirroring the Codex path). Both are off by default and never touch the active/default CLI config; they are backed by theOMNIROUTE_AUTO_SYNC_CODEX_PROFILES/OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILESfeature flags (DB/dashboard override > env > default "false") and additionally gated behind the existingCLI_ALLOW_CONFIG_WRITESwrite-guard. A "CLI profile auto-sync" card on the CLI Code dashboard toggles each (moved from the providers dashboard in #5778 — thanks @rdself). Regression guards:tests/unit/claude-profile-auto-sync-gate.test.ts,tests/unit/codex-profile-auto-sync-gate.test.ts,tests/unit/cli/setup-claude.test.ts(follow-up to #5737). -
cli (startup banner): the
servestartup banner now prints the running OmniRoute version (v3.8.x) beneath the ASCII logo, so the active version is visible at a glance without a separate--versioncall. Regression guard:tests/unit/cli-serve-version-banner.test.ts. Thanks @chirag127 (#5752). -
analytics (subscription cost): flat-rate providers now show $0 in cost analytics instead of an inflated per-token estimate. Subscription / coding-plan providers (every cookie-web provider — ChatGPT Web, grok-web, … — plus the dedicated Minimax Coding, Kimi Coding, GLM Coding, Alibaba Coding Plan, and Xiaomi MiMo plans) bill a flat fee, not per token, yet still carry per-token pricing rows used for estimates — so the analytics dashboard over-reported their cost. A new flat-rate classifier (
src/lib/usage/flatRateProviders.ts) is consulted by the analytics surfaces (analytics route, usage stats, usage analytics) via an opt-inflatRateAsZerocost option, so those providers read $0 while budget / quota / routing keep estimating unchanged. Deliberately NOT zeroed:codex/cx(OmniRoute actively tracks Codex token cost — Fast-tier multipliers, GPT-5.x pricing — and Codex can be a metered account),byteplus(metered ModelArk),minimax-cn(metered China API). Regression guard:tests/unit/flat-rate-cost-5552.test.ts. (#5552) -
mcp (RTK): expose the RTK tool-output learn/discover workflow as two new MCP tools so an agent can grow the RTK filter catalog without leaving the protocol.
omniroute_rtk_discoveranalyzes recently captured raw tool output (discoverRepeatedNoise/suggestFilter) and returns candidate noise patterns plus a suggested filter;omniroute_rtk_learnlists the captured command samples (listRtkCommandSamples) and resolves a command to its RTK filter id (commandToId). Both are read-only (scoperead:compression), wrap the existing RTK discovery primitives (no new logic in the engine), and log to the MCP audit trail. Regression guard:tests/unit/compression/rtk-mcp-tools.test.ts(4). gaps v3.8.42 — T07. -
compression (LLM tier): add an opt-in, default-off LLM-tier compression engine (
llm) that condenses the prose of non-system messages via a pluggable chat-completion backend. It mirrors thellmlinguaengine's contract but is safe by construction: the default backend is a no-op pass-through (the engine never mutates the payload until an operator both enables it and wires a real backend viasetLlmCompressorBackend()), it is not part of the default stacked pipeline,enableddefaults tofalse, fenced code blocks andsystemmessages are never sent to the model, and every backend error fails open (the original segment/body is kept, never thrown). AminTokensfloor skips small prompts. The real production backend is intentionally a VPS-validated follow-up (Hard Rule #18), exactly as thellmlinguaworker backend is gated. Newopen-sse/services/compression/engines/llm/index.ts. Regression guard:tests/unit/compression/llm-compressor-engine.test.ts(8). gaps v3.8.42 — T05/C3. -
memory (typed decay): add opt-in typed memory decay (TV6) so the conversational memory store stops accumulating stale
episodicnoise. Each injected memory now tracks anaccess_count+last_accessed_at(always-on, non-destructive telemetry; migration111_memory_typed_decay), and an opt-in, default-off sweep (MEMORY_TYPED_DECAY_ENABLED, defaultfalse) deletes memories that are past a per-type TTL and not immune. Onlyepisodicdecays by default (30d, env-tunable);factual/procedural/semanticare immune, and any memory accessed>= 3times earns access immunity (mirroring "guardrail/convention/decision never decay"). The decay clock re-bases on the last access, so used memories survive. Deletions reusedeleteMemory(SQLite + sqlite-vec + Qdrant stay in sync) and fail open; an optional periodic sweep is doubly opt-in (also needsMEMORY_TYPED_DECAY_SWEEP_INTERVAL>0). With the flag off nothing is ever deleted (Rule #20 spirit). Newsrc/lib/memory/typedDecay.ts. Regression guard:tests/unit/memory/typed-decay.test.ts(15). gaps v3.8.42 — T10/TV6. -
dashboard (combos): the named-combos editor now lets you drag to reorder the stacked-compression pipeline instead of only editing fixed-position steps. A new pure model (
src/shared/components/compression/compressionPipelineModel.ts) owns add/remove/move/update with the engine→intensity invariant and a never-empty guarantee, and a@dnd-kit/sortableeditor (CompressionPipelineEditor.tsx, matching the sidebar reorder pattern) replaces the inline list inCompressionCombosPageClient. Order persists through the existing combos endpoint. Regression guards:tests/unit/compression-pipeline-model.test.ts(11) +tests/unit/ui/compression-pipeline-editor.test.tsx(4). A dedicatedtests/e2e/compression-studio.spec.ts(Tela A render + tab switch) closes the studios e2e gap the combo-live spec did not cover. gaps v3.8.42 — T06 + T03. -
compression (pipeline): add an opt-in, default-off per-engine circuit-breaker to the stacked compression pipeline (T02). When an engine throws repeatedly across requests, its breaker opens and the stacked loops skip that engine (keeping the body verbatim for that step — fail-open) for a cooldown, then probe once (lazy half-open); success closes it, a failed probe re-opens it. This is distinct from the provider circuit-breaker (
src/shared/utils/circuitBreaker.ts, provider-scoped + DB-persisted) — the newpipelineEngineBreaker.tsis engine-scoped, process-local, and adds zero DB/IO on the hot path. It composes with the existing per-request TV1 bail-out (which skips within a single request); the breaker adds cross-request memory. Default off (COMPRESSION_PIPELINE_BREAKER_ENABLED=false) → byte-identical to the pre-breaker pipeline (a throwing engine still propagates unless TV1 is separately enabled). Configurable per-call, per-CompressionConfig, or via env (_THRESHOLD/_COOLDOWN_MS). Regression guard:tests/unit/compression/pipeline-circuit-breaker.test.ts(9, incl. a throwing-engine integration); existing strategySelector/bail-out suites stay green. gaps v3.8.42 — T02 (2.2). -
compression (CCR): the CCR retrieval-feedback (H8) is now graduated instead of a binary cliff. Previously a block retrieved
>= 3times was flagged do-not-compress and everything below that stayed fully compressible. Now each prior retrieval raises a block's effectiveminCharslinearly (`effectiveMinCh...