Skip to content

Releases: diegosouzapw/OmniRoute

v3.8.51

Choose a tag to compare

@diegosouzapw diegosouzapw released this 30 Sep 01:41
c1e30b7

OmniRoute v3.8.51

2,022 documented changes — 235 features, 1,454 fixes and 333 maintenance entries from 1,973 commits in the cycle, with 301 external contributors. Thank you all.

📄 Complete release notes (every entry, full description and credit): attached below as CHANGELOG-v3.8.51.md, and in CHANGELOG.md → [3.8.51]. GitHub caps a release body at 125,000 characters, so the 1,454 fixes and 333 maintenance entries live there in full.

📊 Release by the numbers

👥 People who contributed 325
📝 Commits in the cycle 1,972
🔀 Pull requests referenced 1,942
📋 Changelog entries 2,022
🙌 Contributors credited in entries 322
🤖 Automated dependency commits 33

Entries by type

Type Count
🐛 Fixes 1443
✨ Features 228
🧹 Chore 135
🧪 Tests 90
📚 Docs 72
⚙️ CI 11
♻️ Refactor 11
⚡ Performance 10
🏗️ Build 8
🔒 Security 7
📦 Dependencies 5
⏪ Reverts 2

🏆 Top 25 contributors this cycle

By commits in 091589089c..4aed4b4a08, author identities consolidated via .mailmap and the merged PR's GitHub login. Bots excluded.

# Contributor Commits
🥇 diegosouzapw 589
🥈 Dizzle (@maxmad64bis) 175
🥉 Bob.Hou (@HouMinXi) 167
4 Paco Cartones (@pacocartones) 93
5 Koosha Paridehpour (@KooshaPari) 50
6 Ravi Tharuma (@RaviTharuma) 43
7 Markus Hartung (@hartmark) 33
8 Nguyen Thanh Dat (@ntdatt812) 31
9 QuangBlue 31
10 Webman (@jonlwheat2-gif) 23
11 Goni Sulaiman (@gonisulaimann) 22
12 lorenzozane (@lorenzozanee) 22
13 anhtahaylove 19
14 backryun 18
15 Fouad Salkini (@fouadSalkini) 16
16 Abhishek Sharma (@abhisheksharma2411) 15
17 Nguyen Thanh Dat (@datrixlab) 15
18 Patryk Kopyciński (@patrykkopycinski) 15
19 Rafa Martins (@rafacpti23) 15
20 Caio Figueiroa (@shipsfromrio) 15
21 initguru 12
22 Jan Leon (@JxnLexn) 12
23 Syed Raheemuddin (@raheemuddin786) 12
24 Xmon Dai (@xiechimon) 12
25 Nguyễn Viết Tuấn (@TheDemonTuan) 11

Highlights

  • Providers and catalogs (42 provider features): new gateways and web-session providers (UC / UC Direct, Lyceum, regolo /v1), GPT-6 Astra/Sol/Luna in the Codex and OpenAI catalogs with effort aliases, account-live listings as the chat catalog source for Claude/Codex/Copilot/AGY, safe Codex model discovery.
  • Routing and resilience: hierarchical concurrency admission, adaptive reasoning effort (auto), expiry-first account fallback, per-kind breaker cooldown escalation, provider-breaker close on a successful probe, local-cooldown 429s no longer read as quota exhaustion.
  • Proxies: pools stop re-serving a member the provider just refused, refused-egress set-aside tuning, egress-IP visibility per pool, selector steering away from recently refused members, persisted blocked-verdict history.
  • Video / Audio Bridge: hardened drill-down cache substrate, bounded transcript contract, opt-in audio extraction + STT, earned "embedded" transcript provenance.
  • Free tier and Radar: eligibility-gated free-tier bucket, customModels[].isFree, free-tier capability in the plugin manifest, Radar rate-limit and rules surfacing.
  • Dashboard and observability: continuous call-log export, orchestration "Repeat", managed-lease sessions view, per-attempt upstream timing on proxy logs (opt-in), logs-export truncation notice.
  • Security hardening (late in the cycle): pre-request hooks in an isolated realm (GHSA-9p9m), IP allow/deny judged on the connection, trailing-dot internal hosts, management auth on OAuth routes, OIDC/Trae/GHE Copilot guards, constant-time token comparison, loopback-only version manager, traversal refused in the Codex /v1/responses subpath, linear-time tool-markup scanners.
  • Release engineering: npm Trusted Publishing (OIDC), lockfile valid for npm 10/11/12 again, AI-attribution gate with a reviewed historical allowlist.

✨ New Features

  • feat(audio): proxy native ElevenLabs voices, text-to-speech, and speech-to-text HTTP routes through stored OmniRoute credentials, preserving query strings, multipart uploads, binary responses, and upstream errors (#10556 #11434) — thanks @hartmark
  • Added Google AI Studio Gemini batch text-to-speech support through POST /v1/audio/speech (#11434) — thanks @hartmark
  • Run synchronous RTK and Caveman request compression in a bounded worker-thread pool, keeping large /v1/responses compression heaps outside the HTTP isolate while preserving strict fail-open behavior and per-engine telemetry (#11434) — thanks @hartmark
  • feat(routing): subscription-first auto groupings — auto/subscription routes only through plan-included connections with a documented hard-stop overage and fails closed on exhaustion, while auto/thrifty orders the pool subscription → keyless → free → cheap → premium and steps up one rung at a time as each is exhausted. Billing class comes from a curated per-connection catalog (uncurated is treated as metered, never plan-included), both reuse STRICT_ZERO_COST's per-connection verification, and a quota reading whose resetAt has passed is now refreshed regardless of TTL so routing returns to plan… (#11146) — thanks @yourspraveen
  • feat(providers): publish a management-authenticated versioned web-session credential contract from OmniRoute's canonical browser credential metadata (#11340) — thanks @Zartharas
  • feat(video bridge): harden the optional drill-down cache substrate with exact-path broker policy, canonical principal/session/media isolation, independent retained-byte quotas, cancellation-safe commits, rejection of excess or non-canonical Base64 padding and non-JPEG/truncated media, warning-sensitive full JPEG canonicalization that strips trailing polyglot bytes, server-derived dimensions, and auditable derivation metadata; production tenant binding and multi-resolution selection remain follow-up work (#11369 #11434) — thanks @hartmark
  • feat(search): Add Xquik X search with typed results, credential validation, REST routing, and MCP selection (#11370) — thanks @kriptoburak
  • feat(video): add an opt-in focused analysis mode that safely uses a normalized, 500-code-point latest-user hint for task-aware frame captions while preserving full-mode prompts, temporal-window isolation, and cache identity without storing raw task text (#11383 #11434) — thanks @hartmark
  • feat(dashboard): surface durable exclusive managed leases in the existing Sessions view, keeping leased clients visible across idle gaps while marking connections with in-flight work as active (#11389) — thanks @KaspaPulse
  • build(bun): allow Turbopack bundler flag on Bun 1.4+ with configurable Webpack fallback (#11471) — thanks @TheDemonTuan
  • feat(api): add an opt-in modelVisibilityAllowlist/modelVisibilityDenylist settings pair to curate exactly which models /v1/models advertises, mirrored into every auto/* combo candidate pool so a denied model cannot be routed to via combo selection either (#11481 #11997)
  • feat(rankings): the Free Provider Rankings page shows what each provider actually served over the last 24 h. It ranked by ELO alone, which left a provider that answers every call with an error in first place; the usage data was already served by the API but never requested. A provider with too small a sample shows a dash, not a number (#11546 #11553) — thanks @maxmad64bis
  • feat(guardrails): enforce a bounded, deterministic contract for Video Bridge transcripts — 256 cues, 4096 input code units and 4 KiB UTF-8 per cue, 64 KiB total text, malformed-Unicode rejection, focus-window scoping, cross-source reconciliation with contributing-source metadata, and a structural provenance trust boundary so caller JSON can never self-assert embedded/audio-bridge provenance (#11652 #12009)
  • feat(video): orchestrate optional Video Bridge audio extraction and Audio Bridge STT behind a dual opt-in (operator setting AND per-request signal) — a new loopback-only broker mode=audio operation shares the frame path's exact process queue, deadline, AbortSignal, and byte budgets to extract a bounded mono 16 kHz PCM WAV from the same already-downloaded video, then reuses the existing Audio Bridge transcription boundary; provider segment timing is preserved when available and marked coarse otherwise, and every failure degrades to a visual-only-safe partial instead of throwing (#11654 #12012)
  • Add a tenant-bound Video Bridge drill-down lifecycle on top of the existing secure cache substrate: opaque hashed handles (never raw session/video identifiers), preview/standard/detail multiresolution variants resampled on read, response pagination capped at 8 frames and 32 MiB, and a new authenticated /api/v1/video-bridge/drilldown consumer route that stays disabled for remote access by default and denies cross-key access with the same response as a nonexistent handle (no existence oracle) (#12006)
  • test(video): Add the Video Bridge FU-07/FU-09 promotion-evidence harness — a frozen Zod manifest schema covering the 8 required scenario kinds (static scenes, rapid cuts, late facts, fades, blur, small text, close events, visual prompt injection) with a minimum of 3 repetitions per case, deterministic declarative fixture recipes (videoBridgePromotionFixtures.ts), a pure medians/p95 metrics aggregator, a pure FU-07/FU-09 promotion-verdict evaluator applying the ticket's exact thresholds (missing token usage always holds), a digest-only persistence layer that never retains raw media or raw model… (#11656 #12008)
  • feat(video bridge): "embedded" transcript provenance can now be legitimately earned instead of merely asserted — a bounded, ...
Read more

v3.8.50

Choose a tag to compare

@diegosouzapw diegosouzapw released this 26 Aug 19:30
5458026

📊 Release by the numbers

👥 People who contributed 248
📝 Commits in the cycle 1,714
🔀 Pull requests referenced 1,666
📋 Changelog entries 1,182
🙌 Contributors credited in entries 256
🤖 Automated dependency commits 22

Entries by type

Type Count
🐛 Fixes 779
✨ Features 169
📚 Docs 29
🧹 Chore 27
🧪 Tests 15
♻️ Refactor 5
⚡ Performance 3
providers 2
🔒 Security 2
⚙️ CI 2
deps 2
maint 2

🏆 Top 25 contributors this cycle

By commits in ed2db6cb19..v3.8.50, author identities consolidated via .mailmap. Bots excluded.

# Contributor Commits
🥇 diegosouzapw 738
🥈 backryun 88
🥉 Dizzle 66
4 Ravi Tharuma 52
5 Markus Hartung 48
6 Bob.Hou 42
7 Rouzbeh† 38
8 Xiangzhe 31
9 Paco Cartones 28
10 Nguyen Thanh Dat 23
11 Aman 22
12 Will Gordon 19
13 小妍儿 ✨ 17
14 adevwithpurpose 16
15 Andrew B. 10
16 NOXX - Commiter 10
17 ignamiranda 10
18 Jonathan Bailey 9
19 Ke Jin 9
20 Austin Liu 8
21 Chewji 8
22 Prudhvi Vuda 7
23 benzntech 7
24 rinseaid 7
25 stanley 7

Living section — regenerated 2026-08-12 from all cycle commits (cycle open ed2db6cb19 → tip). Bullets carry the merged PR and its author; direct pushes listed separately.

✨ New Features

  • feat(search): first-class X Search provider (x-search) on… #10985

  • feat(core): add Layer A capability filter at router #5696

  • feat(providers): add DeepAI as paid API-key image provider #6671

  • feat(providers): add Naga.ac and… #6674 — thanks @chirag127

  • feat(api): add response content encoding verification —… #6736

  • feat(api): add plugins marketplace install endpoint with… #6752

  • feat(chatgpt-web): harden prompt-emulated… #7679 — thanks @horacecar

  • docs: add management authentication terminology guide #7786

  • feat(a2a): Conductor bridge — long-lived SSE consumer that… #8080

  • feat(a2a): the Agent Card (/.well-known/agent.json) now… #8119

  • feat(dashboard): "Conductor" panel — OmniConductor fleet… #8221

  • feat(dashboard): Faro chat with voice on the Conductor panel —… #8222

  • feat(a2a): inbound delegation to the OmniConductor fleet — `POST… #8223

  • docs: add low-memory/small VPS optimization guide #8237

  • feat(providers): add connection-level… #8369 — thanks @Benson-mk

  • feat(copilot): add approval gate for runOmniRouteCli commands #8461

  • feat(ci): add windows-latest leg to test-bun-sqlite job #8468

  • feat(electron): Desktop app can now attach… #8799 — thanks @soulhakr

  • Database The node:sqlite fallback now… #8870 — thanks @artickc

  • feat(models): add exact per-model… #8908 — thanks @xz-dev

  • Providers expands the Novita AI catalog… #8913 — thanks @jax-novita

  • feat(providers): native xAI Agent Tools passthrough on… #8964

  • feat(providers): add UnoRouter provider UnoRouter is an… #8978

  • feat(sse): deprecated the legacy gemini-cli **upstream… #7034 #8980

  • Add a default-off connection setting for Codex, OpenAI, and…

  • Omit opaque encrypted reasoning values from persisted call logs… #9000

  • feat(providers): add Regolo AI OpenAI-compatible provider #9031

  • feat(db): add provider-scoped model aliases that survive… #9068

  • feat(cursor): surface a dismissible dashboard banner suggesting… #9173

  • feat(cursor): proactively renew Cursor sessions before their ~24h… #9173

  • feat(codex): accept parenthesized… #9208 — thanks @seakleangnhak

  • feat(usage): surface Claude thinking token… #9214 — thanks @luoyide

  • feat(ollama): add Ollama Local embedding… #9225 — thanks @HaoNgo232

  • feat(images): execute full combo strategy + fallback in… #9239

    Adds open-sse/services/imageCombo.ts that expands combo targets, filters to images-capable, executes the priority strategy with handleImageGeneration per target, and returns the first success or last failure. Route patches detect combo names before model resolution and divert to the new execution path.

  • feat: make forwarded upstream response-header budget configurable… #9243

  • feat(providers): filter provider detail… #9247 — thanks @RobertsXML

  • feat(providers): make video_url… #9248 — thanks @HellFiveOsborn

  • feat(gemini): recursive type:object injection in schema… #9268

  • feat(dashboard): render a conditional "Get API key" link on the… #9270

  • feat(providers): accept JSON cookie… #9284 — thanks @AIB1TAL0S

  • feat(providers): support max reasoning effort for opencode-zen… #9318

  • feat(providers): expanded the NanoGPT (nano-gpt.com) upstream… #9322

  • feat(sse): combo system_message supports server-side… #5501

  • feat(sse): template expansion… #5501 #9414 — thanks @maxmad64bis

  • feat(sse): New-API/One-API/Sub2API aggregator balance detection… #9415

  • feat(catalog): added opt-in settings hideAutoCombos and… #9418

  • feat(opencode-plugin): added… #9473 — thanks @omniroute

  • feat(providers): add native DeepSeek V4 Flash and Pro… #9485

  • feat(opencode-plugin): warm catalog startup from disk snapshot +… #9490

    The config-shim hook now reads the last disk snapshot before fetching, so the provider registers immediately with the last-known-good catalog (~1-2s vs ~30s on a warm gateway). All six fetchers run concurrently via Promise.allSettled instead of sequentially. A failed refresh keeps the snapshot (no overwrite). An in-flight guard prevents concurrent refreshes for the same cache key. The features.diskCache: false opt-out disables the warm read entirely.

  • feat(models): Test All's "Auto-hide failed models" no longer… #9511

  • Add an advisory forgotten-sibling-tests report to pull-request… #9530

  • feat(providers): add Muse Code CLI provider preset #9544

  • feat(plugins): expose client request headers in plugin… #9570

  • feat(plugins): add onStreamComplete built-in event exposing… #9571

    Adds a new onStreamComplete plugin event that fires after an SSE stream is fully
    consumed, carrying usage token counts and timing metrics (latency, TTFT). Built-in
    events now include onStreamComplete as a fire-and-forget lifecycle hook.

    Payload: status, usage (prompt_tokens, completion_tokens, reasoning_tokens,
    cache_read_input_tokens, cache_creation_input_tokens), timing (latencyMs, ttft),
    model, provider, errorCode.

    Non-breaking — existing onResponse hooks with { streamed: true } remain unchanged.

  • feat(audio): Soniox STT + TTS provider (sx) — async… #9579

  • Show cache-read and cache-write token counts in request log rows and…
    report them. (#9620)

  • feat(memory): support custom OpenAI-compatible endpoints for… #9622

  • feat(resilience): add an opt-in watchdog for persistently slow… #9709

  • Onboarding: add an explicit, reviewable one-click setup for eligible…
    with per-provider caution links, selectable confirmation, idempotent creation, and safe partial
    retries. Existing provider connections are never changed and setup completion never enables
    providers silently. (#9752)

  • feat(settings): add a dedicated Modality Bridge settings page… #9782

  • feat(modality bridge): Transcribe chat audio for text-only… #9807

  • feat(memory): PROVIDERS_SYSTEM_MUST_BE_FIRST (the… #6135 #7293 #9924

  • feat(api): API keys can disable prompt… #10001 — thanks @shixi-li

  • Add cliproxy provider exposure controls and manifest injection #7329 — thanks @KooshaPari

  • feat(infra): add a systemd autostart unit for Linux #8635

  • feat(db): add node sqlite adapter parity #8871 — thanks @epsilonode

  • feat(alibaba): free-tier routing with live quota sync #8893 — thanks @AndrianBalanescu

  • feat(oauth): add Raycast Pro provider with local auto-import #8895 — thanks @AndrianBalanescu

  • feat(executors): add isolated Claude Code bridge over Devin ACP #8914 — thanks @McLuck

  • feat: improve provider quota layouts #8916 — thanks @apoapostolov

  • feat(mcp): add omniroute_create_combo tool #8925 — thanks @lucasmellos

  • feat(ci): gate the publish on clean-install AND upgrade-over-previous #8953

  • feat(providers): add Conol (conol.ai) web session provider #8974 — thanks @artickc

  • Feat/combo provider wise model test #9011 — thanks @JoshimOfficial

  • feat(model-alias): add runtime Model Alias Resolver middleware #9020 — thanks @Egorich-print

  • feat(i18n): complete zh-CN localization for compression engines and dashboard UI #9038 — thanks @qianze0628

  • feat: Cheaper Inference provider (chat + native Responses + images, sponsor rail 2nd) #9043

  • feat(providers): add comprehensive support for self-hosted Firecrawl via FIRECRAWL_BASE_URL and custom base URLs #9052 — thanks @mad-gooze

  • feat(dahl): add manual API key option alongside auto-generated token #9077 — thanks @pizzav-xyz

  • feat(ci): G0 — reforça o trilho PR→release/ ** #9108

  • feat(providers): native xAI Agent Tools passthrough for /v1/responses #9111 — thanks @VXNCXNX

  • feat(.50): completa itens restantes — G13, G14, gap34, docs, R0.2 #9126

  • feat(g1): rewrite combo-strategy check to runtime-import approach #9131

  • feat(test:scoped): TIA-based local test runner (#8084 D1) #9143

  • feat(docker): publish next from active release branches #9181 — thanks @Zartharas

  • feat(usage): show Grok Build billing limits #9205 — thanks @xz-dev

  • feat(models): functional gateway mirrors + fix synced-substitution #9217

  • feat(admission): add adaptive overload protection for LLM routes #9262 — thanks @xz-dev

  • feat(i18n): update italian translations #9280 — thanks @Gecky2102

  • feat(dashboard): persist provider screen filters to URL for bookmarking #9307 — thanks @swingtempo

  • feat(api-manager): add provider-level model permissions #9313 — th...

Read more

Radar catalog export (rolling)

Choose a tag to compare

@github-actions github-actions released this 20 Aug 11:19
0f13fe4

Export estável do catálogo OmniRoute para o Radar. Atualizado automaticamente; NÃO é um release de versão do produto.

v3.8.49

Choose a tag to compare

@diegosouzapw diegosouzapw released this 30 Jul 00:37
c9d4a45

All 1383 entries from this cycle are listed below, one line each — descriptions are
trimmed to fit GitHub's 125,000-character release body. Full wording, context and links:
CHANGELOG.md.
Living section — regenerated 2026-07-19 from all 306 cycle commits (bump 2c62333 → tip). Bullets carry the merged PR and its author; direct pushes listed separately. Finalized at the v3.8.49 release.

✨ New Features

  • feat: generalize ensureThinkingBudget to all providers +… (#6979) — @rafaumeu

  • feat(6922): effort-tier aliases for glm-5.2 & mimo-v2.5 on… (#6987) — @rafaumeu

  • feat(providers): curated OpenRouter embeddings catalog + specialty merge… (#6994)

  • feat(quota): opt-in auto-ping to keep Codex quota windows warm (#6995)

  • feat(providers): add Agnes AI native provider support (#7035) — @HouMinXi

  • feat(sse): allow disabling : comment heartbeats via… (#7036) — @xier2012

  • feat(perf): add performance.mark/measure to SSE pipeline +… (#7045) — @oyi77

  • feat(providers): add Dahl free inference provider (#7062) — @growab

  • feat(ci): boot-smoke the packed npm tarball (check:pack-boot,… (#7086)

  • feat(ci): hotfix fast-lane + tests-only E2E skip (WS3.1) (#7088)

  • feat(ci): continuous release-green — on-push quick gate + 3x/day… (#7089)

  • feat(ci): duration-balanced E2E shards via LPT bin-packing (WS4.1) (#7090)

  • feat(ci): TypeScript 7 native shadow for typecheck:core (WS4.2,… (#7091)

  • feat(release): npm staged publishing + pre-publish boot-smoke (WS1.3) (#7092)

  • feat(release): post-publish verifier — clean-container install + boot… (#7109)

  • feat(ci): Mergify merge queue + manual-train fallback runbook… (#7112)

  • feat(ci): Windows leg for Electron prepare smoke (WS1.5) (#7113)

  • feat(ci): Codecov patch coverage (informational) + fix missing… (#7114)

  • feat(sidecar): support conditional provider manifest refresh (#7130) — @KooshaPari

  • feat(homolog): real-environment E2E homologation suite (npm run… (#7133)

  • feat(usage): add Codex reset credit picker (#7154) — @JxnLexn

  • feat(ci): Trunk Flaky Tests uploads for vitest + Playwright E2E… (#7175)

  • feat(ci): Trunk Flaky Tests upload on the fast-path vitest job… (#7205)

  • feat(kiro): register GPT-5.6 Sol/Terra/Luna model family (#7209)

  • feat(dashboard): show Codex plan label in provider and quota views (#7210)

  • feat(dashboard): add reorder connections by availability button (#7211)

  • feat(dashboard): add 180D and 365D usage/cost analytics periods (#7213)

  • feat(api): add Vary: Accept-Encoding to token-authenticated /v1*… (#7217)

  • feat(api): expose GET /api/usage/model-latency-stats (#7218)

  • feat(dashboard): add compression-mode selector to Context & Cache combos… (#7219)

  • feat(sse): route GitHub Copilot Claude models through native… (#7223)

  • feat(mitm): add Antigravity reasoning-effort overrides (#7228)

  • feat: replace free-text model inputs with hidePaid-aware… (#7229)

  • feat: editable ComfyUI base-URL field + per-connection… (#7232)

  • feat(sse): add optional-enum null-omission idiom for codex… (#7233)

  • feat(sse): preserve tools/tool_choice for tool-bearing requests… (#7235)

  • feat(api): accept x-goog-api-key header for client-facing auth (#7236)

  • feat(sse): add native xAI Grok Imagine video generation provider (#7238)

  • feat: add Type filter and easiest-first sort to Free Provider… (#7240)

  • feat(cli): add Grok Build CLI tool setup (~/.grok/config.toml) (#7241)

  • feat(provider): add Chenzk API OpenAI-compatible gateway (#7246)

  • feat(providers): let custom connections opt into prompt-cache capability (#7257)

  • feat(db): include xp_audit_log in automatic retention/prune (#7260)

  • feat(api): structured X-Routing-Fallback-Reason header for relay… (#7262)

  • feat(compression): support RTK TOML schema v1 filters (#7281) — @JxnLexn

  • feat: add principal-scoped CCR MCP lifecycle (#7282) — @JxnLexn

  • feat(morph): refresh curated models (#7314) — @backryun

  • feat(issue-agent): surface RecordedTriageTimeoutError as 504 (#7315) — @KooshaPari

  • feat(incident-response): structured incident response templates (#7334) — @KooshaPari

  • feat(providers): add xAI OAuth PKCE provider (#7399) — @fenix007

  • feat(models): advertise Claude reasoning-effort variants in /v1/models (#7497) — @thepigdestroyer

  • feat(kimi): sync Code, Web, and Moonshot providers (#7531) — @backryun

  • feat(resilience): guard OmniRoute peer routing loops (#7555) — @isiahw1

  • feat: add Mixedbread AI as embeddings provider (#7595)

  • feat(providers): add Rev AI speech-to-text provider (#7596)

  • feat: add Freepik (Magnific Mystic) image generation provider (#7597)

  • feat(sse): add DeepInfra as a video-generation provider (#7598)

  • feat(providers): add Felo chat-aggregator provider (#7599)

  • feat(sse): add Notion AI Web (Unofficial/Experimental) provider (#7600)

  • feat: add FreeTheAi as OpenAI-compatible gateway provider (#7602)

  • feat: add Gladia as an async speech-to-text provider (#7603)

  • feat: add EdgeTTS audio-tts provider (#7605)

  • feat(video): add Novita AI as video-generation provider (#7606)

  • feat: add Segmind image+video provider (#7608)

  • feat: add Microsoft Designer as image provider (#7609)

  • feat: per-model default reasoning_effort + no-think none on… (#7631)

  • feat(sse): per-model upstream header-response timeout override (#7632)

  • feat(dashboard): in-product guidance for prompt compression engines (#7634)

  • feat(usage): add TTFT/E2E-latency/tokens-per-second to model latency… (#7635)

  • feat: import providers from CSV/JSON file (#7636)

  • feat: confirm before removing a single connection (#7640)

  • feat(sse): honor excluded models in no-auth auto-combo candidate… (#7646)

  • feat(providers): add g4f.space no-key gateway… (#7647)

  • feat: rate-limit queue admission control (maxQueueDepth + 15s… (#7649)

  • feat(sse): generalize session affinity TTL to all providers (#7650)

  • feat: OpenRouter quota tracking (key/credits + free-window… (#7651)

  • feat(sse): quota tracking for AgentRouter, v0 (Vercel), FreeModel… (#7653)

  • feat(providers): Speechmatics STT, gTTS, VibeProxy preset (#6659, #6667,… (#7655)

  • feat(api): route Google AI Studio Imagen through… (#7656) — @danscMax

  • feat(auth): OIDC as optional dashboard admin login gate (password… (#6973) — @mikolaj92

  • feat(api): add pagination params to 8 DB modules + recharts… (#7046) — @oyi77

  • feat(proxy): operator-level proxy subscriptions (Karing-style) —… (#7299) — @xier2012

  • feat(grok-cli): align with official Grok Build client (#7358) — @backryun

  • feat(providers): Complete GHE Copilot OAuth provider implementation (#7546) — @hppsc1215

  • feat(guardrails): add CredentialMaskerGuardrail for API key/secret… (#7683) — @Securiteru

  • feat(perplexity): refresh provider integrations (#7687) — @backryun

  • feat(providers): notion-web live model discovery via getAvailableModels (#7696) — @artickc

  • feat(providers): add proactive cf_clearance/User-Agent hint to grok-web… (#7713)

  • feat: add live gRPC-web quota fetcher for grok-cli (#7714)

  • feat(api): add opt-in auto-sync scheduler for free-proxy sources (#7716)

  • feat(dashboard): show proxy name in badge, sort saved-proxy picker,… (#7720)

  • feat(cli): add auth export command for decrypted provider… (#7724)

  • feat(oauth): accept full ChatGPT session JSON for Codex manual import (#7725)

  • feat(sse): add nvidia NIM local RPM budget + concurrency cap (#7726)

  • feat(gemini-web): emulate OpenAI tool calling via the webTools prompt shim (#7727)

  • feat(services): introduce pluggable service-provider contract, migrate… (#7730)

  • feat(mitm): root-CA + per-host leaf certs for AgentBridge static… (#7731)

  • feat(providers): add hailuo-web (MiniMax web) chat provider (#7734)

  • feat: browser login for Grok Build provider (#7735)

  • feat(routing): wire interceptFetch tool interception into the chat… (#7736)

  • feat(sse): add X-OmniRoute-Decision routing trace header (#7765)

  • feat(providers): zai-web live model discovery with local-catalog fallback (#7766)

  • feat(api): sync upstream reasoning.supported_efforts into… (#7767)

  • feat(dashboard): pin Kimi providers first in category + official… (#7775)

  • feat(chaos+ponytail): parallel chaos-mode dispatch + ponytail output … (#7781) — @Moseyuh333

  • feat(perf): IC2 — cache provider connections by ID + lazy-decrypt… (#7787) — @oyi77

  • feat(quality): gate the free-tier headline so it can never silently… (#7798)

  • feat(providers): expose an explicit tier override for any provider… (#7838)

  • feat(routing): read-only auto/* candidate transparency + per-API-key… (#7839)

  • feat(catalog): map unmapped free tiers, add navy + aihorde, surface… (#7840)

  • feat(providers): add OpenRouter speech-to-text (audio transcription)… (#7861) — @Tasogarre

  • feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] (#7874) — @backryun

  • feat: support Bun bundled SQLite runtime (#7878) — @Arul-

  • feat(providers): add 5 free-tier providers (ainative, aion, sealion,… (#7887)

  • feat(vnc-session): persistent noVNC browser login for web-cookie providers (#7892) — @Capslockb

  • feat(sse): add PromptQL playground provider (unofficial) (#7911) — @artickc

  • feat(cline): align ClinePass catalog and request protocol (#7914) — @backryun

  • feat: narrow mcp:connect scope + per-key HTTP tool-scope… (#7967)

  • feat: provider tab account search + mirrored top pagination (#7968)

  • feat: canonical numeric helpers + tier-1 (analytics) migration (#7969)

  • feat(sse): add HyperAgent (hyperagent.com) unofficial web provider (#7994) — @artickc

  • feat: copilot-m365...

Read more

v3.8.48

Choose a tag to compare

@diegosouzapw diegosouzapw released this 13 Jul 21:19
7ee5bbc

⚠️ Hotfix release. The published npm package for 3.8.47 crashed on every boot (#7065) and was deprecated — 3.8.48 is the first installable release of the v3.8.47 cycle, so everything listed under [3.8.47] below ships here.

🐛 Bug Fixes

  • fix(build): ship dist/head-response-guard.cjs in the npm tarball — the prepublish prune allowlist lacked it, so every omniroute boot of the published 3.8.47 crashed with ERR_MODULE_NOT_FOUND (3rd occurrence of this class after tls-options/3.8.41); now allowlisted, enforced by check:pack-artifact, and guarded by a closure test that derives every server-ws.mjs sibling import (#7065, #7040)
  • fix(build): Electron Windows packaging — the better-sqlite3 Electron-ABI rebuild now spawns npx.cmd through a shell (Node's CVE-2024-27980 hardening made the shell-less spawn fail with status null on Windows runners, breaking the v3.8.47 desktop build)
  • fix(ci): Sonar quality gate zeroed on new code — the coverage lcov now reaches the scanner at coverage/lcov.info (it read 0% on every scan), the async isCloudEnabled() gate in the Kiro auto-import route is awaited (cloud sync ran even when disabled), the dead structuredClone fallback in the reasoning-split clone is a real JSON fallback, the codex executor handles the async reader.cancel() rejection, deterministic localeCompare sorts, a path-traversal guard in classify-pr-changes.mjs, and the Docker better-sqlite3 rebuild uses npm's bundled node-gyp instead of npx --yes
  • chore(ci): the Sonar quality gate is informational (sonar.qualitygate.wait=false) while the org's SonarCloud plan cannot associate the tuned "OmniRoute way" gate (coverage ≥60 aligned with the repo floor)

📦 Everything from the v3.8.47 cycle ships here

The 3.8.47 npm package was never installable (#7065), so 3.8.48 is the release that actually delivers the whole v3.8.47 cycle — full notes below:

  • 9router Codex import: the Codex bulk-import endpoint (POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt + nested providerSpecificData), not just snake_case — normalizeCodexImportRecord maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard: tests/unit/codexBulkImport.test.ts (9router camelCase record, pre-supplied providerSpecificData without an id_token, snake_case-not-overridden, and a full {accounts:[...]} flatten).

✨ New Features

  • feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)

  • feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)

  • feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)

  • feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)

  • feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)

  • feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)

  • feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint. (#6709 — thanks @diegosouzapw)

  • feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)

  • feat(release): changelog.d/ fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw)

  • feat(quality): validate-release-green --full-ci reproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw)

  • feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existing common.search/common.noResults i18n keys (zero new locale edits) and the shared Input icon="search" pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted into filterSidebarSectionsByQuery() (src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard: tests/unit/sidebar-search-filter.test.ts, src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber)

  • feat(combo): auto/* combos gain a strict budget-cap fallback policy — X-OmniRoute-Budget-Fallback: strict (or the persisted config.budgetFallback: "strict") makes an over-budget request fail fast with HTTP 402 instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existing X-OmniRoute-Budget/X-OmniRoute-Mode per-request controls (#6023/#6024/#6025), consolidated into resolveRequestAutoControls(). Regression guard: tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470)

  • Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)

  • Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the #3587 reasoning-model max_tokens buffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev)

  • feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (comboStickyRoundRobinLimit, resolved via resolveComboStickyRoundRobinLimit() — per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A new providerStrategies setting (Zod-validated map, src/shared/validation/settingsSchemas.ts) lets a specific provider override the global fallbackStrategy/stickyRoundRobinLimit without touching the account-wide default, wired into getProviderCredentials() (src/sse/services/auth.ts) ahead of the global fallback. Regression guard: tests/unit/combo-rr-sticky-9router.test.ts, tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen)

  • feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG → @lobehub/icons React components → thesvg.org CDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard: tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7)

  • Skill Collector CLI detection: new GET /api/skills/collect/detect + POST /api/skills/collect/install (and the cli-skill-collector agent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally via getCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes and GET/POST /api/github-skills now require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES + SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). The omniroute_github_skills_install MCP tool now reports the honest action: "planned" instead of "installed", matching the REST route (#6294 — thanks @Moseyuh333)

  • ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to cp (matching the OAUTH_PROVIDERS catalog alias) so <alias>/<modelId> routing resolves correctly, the OAuth refresh dispatch now routes clinepass to the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard: tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok)

  • feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no aorAAAAAG-prefixed refresh token) a...

Read more

v3.8.47

Choose a tag to compare

@diegosouzapw diegosouzapw released this 13 Jul 14:57
e8950de
  • 9router Codex import: the Codex bulk-import endpoint (POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt + nested providerSpecificData), not just snake_case — normalizeCodexImportRecord maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard: tests/unit/codexBulkImport.test.ts (9router camelCase record, pre-supplied providerSpecificData without an id_token, snake_case-not-overridden, and a full {accounts:[...]} flatten).

✨ New Features

  • feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)

  • feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)

  • feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)

  • feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)

  • feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)

  • feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)

  • feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint. (#6709 — thanks @diegosouzapw)

  • feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)

  • feat(release): changelog.d/ fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw)

  • feat(quality): validate-release-green --full-ci reproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw)

  • feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existing common.search/common.noResults i18n keys (zero new locale edits) and the shared Input icon="search" pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted into filterSidebarSectionsByQuery() (src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard: tests/unit/sidebar-search-filter.test.ts, src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber)

  • feat(combo): auto/* combos gain a strict budget-cap fallback policy — X-OmniRoute-Budget-Fallback: strict (or the persisted config.budgetFallback: "strict") makes an over-budget request fail fast with HTTP 402 instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existing X-OmniRoute-Budget/X-OmniRoute-Mode per-request controls (#6023/#6024/#6025), consolidated into resolveRequestAutoControls(). Regression guard: tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470)

  • Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)

  • Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the #3587 reasoning-model max_tokens buffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev)

  • feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (comboStickyRoundRobinLimit, resolved via resolveComboStickyRoundRobinLimit() — per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A new providerStrategies setting (Zod-validated map, src/shared/validation/settingsSchemas.ts) lets a specific provider override the global fallbackStrategy/stickyRoundRobinLimit without touching the account-wide default, wired into getProviderCredentials() (src/sse/services/auth.ts) ahead of the global fallback. Regression guard: tests/unit/combo-rr-sticky-9router.test.ts, tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen)

  • feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG → @lobehub/icons React components → thesvg.org CDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard: tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7)

  • Skill Collector CLI detection: new GET /api/skills/collect/detect + POST /api/skills/collect/install (and the cli-skill-collector agent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally via getCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes and GET/POST /api/github-skills now require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES + SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). The omniroute_github_skills_install MCP tool now reports the honest action: "planned" instead of "installed", matching the REST route (#6294 — thanks @Moseyuh333)

  • ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to cp (matching the OAUTH_PROVIDERS catalog alias) so <alias>/<modelId> routing resolves correctly, the OAuth refresh dispatch now routes clinepass to the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard: tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok)

  • feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no aorAAAAAG-prefixed refresh token) and can't refresh through the AWS OIDC/Kiro-social path, so tryAwsSsoCache() now detects them (authMethod/provider === "externalidp") and refreshes via the org IdP's own tokenEndpoint (public-client OAuth2 refresh grant, no client secret), persisting TokenType: EXTERNAL_IDP gating so the runtime executor sends the header the AWS CodeWhisperer API requires for these accounts; tokenEndpoint is SSRF-guarded against an HTTPS + known-IdP-host-suffix allowlist. (#6363 — thanks @artickc)

  • Kiro long-lived API key auth: new /api/oauth/kiro/api-key route + KiroService.validateApiKey let a Kiro account be linked with a long-lived AWS CodeWhisperer/Kiro API key instead of the interactive OAuth device flow, with live per-account model discovery (ListAvailableModels, 5-minute cache) layered over the existing static registry fallback (#6587 — thanks @strangersp)

  • Chaos Mode: multi-model parallel/collaborative task execution — dispatches a task to every active provider connection at once (parallel) or chains outputs sequentially so each model builds on the previous one's answer (collaborative), configurable via Dashboard → Chaos Mode (GET/PUT/DELETE /api/chaos/config) and gated per-API-key via a new chaosModeEnabled permission (opt-in — disabled by default globally and per key). POST /api/chaos/run (dashboard session) and POST /api/skills/collect/chaos (external Bearer-token) delegate to a shared executeChaosRun() engine (src/lib/chaos/chaosExecutor.ts) that dispatches in-process via the established synthetic-Request/route-handler pattern (no network hop, no hardcoded port), with a concurrency cap (max 10 parallel), configurable max_tokens (256–128k), a clear error when stream is requested, and collaborative-chain info (provider order + input size). Fixes external Bearer-auth bypass and stale config-cache leakage. Regression guard: tests/unit/chaos-config.test.ts, tests/unit/chaos-executor.test.ts, tests/unit/chaos-api-routes.test.ts. (#6728 — thanks @Moseyuh333)

  • feat(cli): 2 new...

Read more

v3.8.46

Choose a tag to compare

@diegosouzapw diegosouzapw released this 07 Jul 16:36
92715c8

✨ New Features

  • feat(sse): hide paid-only models from auto/* routing when hidePaidModels is on (#6512) — follow-up to #6328/#6495. PR #6495 hid paid-only models from the GET /v1/models listing, but auto/* combos (auto/best-coding, auto/glm, …) could still pick a paid-only backend into their candidate pool → a 402/403 at request time. createVirtualAutoCombo now filters the candidate pool through the new pure open-sse/services/autoCombo/paidModelFilter.ts (filterPaidOnlyCandidates), applying the same free-model predicate #6495 uses in catalog.ts (providerHasFreeModels(provider) && isFreeModel(provider, {id})) whenever settings.hidePaidModels === true. Applied before the category/tier/family narrowing, so it covers every auto/* combo; an all-paid pool degrades to the existing graceful empty-pool path. Opt-in — default OFF leaves the pool unchanged (identity). Regression guard: tests/unit/autoCombo/paid-model-filter-6512.test.ts (4, incl. the default-off identity guard).
  • feat(sse): provider-family auto combos — auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini (#6453) — new routable ids that materialize an on-demand virtual combo spanning whatever installed backends currently expose that model family, degrading gracefully as backends rotate. A new pure open-sse/services/autoCombo/modelFamily.ts (detectModelFamily) classifies by model-id prefix for six families; zai is instead resolved by provider id (z.ai's hosted API serves the same glm-* model ids as every other GLM backend, so auto/zai means "route to my z.ai backend specifically" vs auto/glm's "any connected GLM backend"). Reuses the existing createVirtualAutoCombo on-demand materialization path (no DB writes) and the /v1/models catalog advertising loop. Regression guard: tests/unit/autoCombo/provider-family-combos.test.ts (11).
  • feat(proxy): native proxy-pool round-robin / egress IP rotation (#6365) — a scope (global / provider / account) can now hold multiple proxies as a pool with a rotation strategy, so outbound requests cycle their egress IP instead of pinning one proxy per scope. Migration 117_proxy_pool_rotation.sql lifts the UNIQUE(scope, scope_id) constraint (rebuild via the canonical rename/copy/drop; existing single assignments become 1-element pools) and adds a proxy_scope_rotation companion table holding the per-scope strategy + a persisted monotonic round-robin cursor. Strategies: round-robin (default, monotonic cursor — never Math.random), random, and sticky-per-N-min. Resolution (resolveProxyForScopeFromRegistry / resolveProxyForConnectionFromRegistry) now fetches the alive, position-ordered candidate set (unchanged PROXY_ALIVE_PREDICATE) and applies the strategy; an empty / all-dead pool still returns null — the #6246 fail-closed guard is untouched (never falls through to direct egress). Backend + DB only; dashboard pool-builder UI is a follow-up. Regression guard: tests/unit/proxy-pool-rotation-6365.test.ts (8, incl. fail-closed + backward-compat).
  • feat(providers): end-to-end tool/function calling on the native Gemini /v1beta endpoint (#6222) — both directions of the Gemini↔OpenAI conversion now preserve tool data (previously silently dropped). Request side: convertGeminiToInternal (extracted to its own testable module) maps tools[].functionDeclarations → OpenAI tools, prior functionCall parts → assistant tool_calls, and functionResponse parts → tool-role messages. Response side: convertOpenAIResponseToGemini emits parts[].functionCall {name,args} from message.tool_calls, and the streaming openAIChunkToGeminiChunk accumulates fragmented tool_calls deltas by index into complete functionCall parts. The non-Gemini client paths (Claude, OpenAI-Responses) already preserved tool calls — this closes the gap specific to the native Gemini surface. Regression guard: tests/unit/v1beta-gemini-tool-calling-6222.test.ts (6, incl. a streaming SSE round-trip).
  • feat(providers): copilot-m365-web enterprise / work tier support (#6334) — mirrors the EDU-tier pattern (#6210): M365ConnectionParams gains an agent field, a new opt-in M365_ENTERPRISE_OVERRIDES preset (agent=work, scenario=officeweb, licenseType=Premium) applies via providerSpecificData.tier="enterprise" (alias "work"), and agent is also overridable directly via providerSpecificData.agent. buildWsUrl was hardcoding agent="web" (the one enterprise-distinguishing param with no override path), so a Premium work account handshook then returned an empty stream. The individual and EDU paths are untouched. Kilo's dup flag vs #6210 (EDU tier) was a false positive — different tier. Regression guard: tests/unit/copilot-m365-enterprise-6334.test.ts (7). End-to-end confirmation on a real Premium work account is a live-VPS validation follow-up (Hard Rule #18). (thanks @Forcerecon)
  • feat(api): standardized, provider-agnostic effort + thinking request params (#6241) — a thin standardization layer over the existing mature per-provider reasoning plumbing (no provider mapper touched). providerChatCompletionSchema gains a canonical effort (reusing the shared none/low/medium/high/xhigh vocabulary — the UI tiers extra/max collapse onto xhigh) and a boolean thinking. A pure normalizeReasoningRequest (wired once in src/sse/handlers/chat.ts, before any reasoning field is read) folds them onto the fields the translators already consume (reasoning_effort / reasoning.effort / thinking), so they fan out to Anthropic / Gemini / xAI / Responses — an explicit client reasoning_effort / object-shaped thinking always wins (backward-compatible). /models additively exposes supportsThinking + effort_tiers so the frontend can render the toggles (UI component is a follow-up). Regression guard: tests/unit/effort-thinking-standardization-6241.test.ts (12). (thanks @Iammilansoni, @shabeer)
  • feat(combo): new pipeline (sequential) combo strategy (#6297) — the 18th routing strategy runs targets in order, threading each step's output into the next step's input, with an optional per-step prompt (system instruction); only the final step's response is returned. Distinct from fusion (parallel fan-out + judge). Implemented as a self-contained open-sse/services/pipeline.ts (sibling to fusion.ts), dispatched from combo.ts; the step list reuses combo.models order and reads an optional prompt off each target (backward-compatible — ignored by every other strategy). Intermediate steps run non-streaming with tools stripped (complete prose to thread forward); the final step keeps the client's stream flag + tools. A failing/empty/unparseable intermediate step fails the whole pipeline explicitly via a sanitized error (never silently swallowed). Kilo's dup flag vs #563 was a false positive (that's model→chain selection; this is a sequential chain). Regression guard: tests/unit/combo-pipeline-strategy.test.ts (5). (thanks @ofekbetzalel)
  • feat(ci): check:test-masking now flags inline-reimplemented prod conditions (#6348) — a new report-only subcheck (v2, 6A.10 family) catches the wrong-shape contract test: a test that recomputes the condition under test inline instead of importing/exercising the real function (the #6216 class, where === 500 → >= 500 stayed green because the test re-implemented the branch). For each added/modified test file it warns when the file textually duplicates a ≥3-token conditional from a production file touched in the same PR and does not import the symbol/module owning it, via a pure, fixture-tested findReimplementedConditions() with an allowlist mirroring assertReductionAllowlist. Report-only for now (does not fail the gate) — to be promoted to blocking after a triage cycle. Regression guard: tests/unit/check-test-masking.test.ts (45).
  • feat(sse): per-connection routing override (native vs CLIProxyAPI) (#6339) — the previously-dead isCliproxyapiDeepModeEnabled helper is now wired into resolveExecutorWithProxy: a single connection can opt itself into the CLIProxyAPI passthrough executor via providerSpecificData.cliproxyapiMode="claude-native", with precedence connection override > provider upstream_proxy_config mode > default. resolveExecutorWithProxy now receives the resolved connection's providerSpecificData (threaded from chatCore.ts), so one connection can deep-route while the provider's other connections stay native — no DB schema change (the toggle rides in providerSpecificData). Also resolves the same-provider mixing ask in #6340. Regression guard: tests/unit/chatcore-executor-proxy.test.ts (9). (thanks @RaviTharuma)
  • feat(dashboard): "Add session cookie" modal now shows a prominent "Open ‹host› →" link to the provider's own site (#6268) — every -web cookie-session provider (chatgpt-web, claude-web, gemini-web, kimi-web, lmarena, qwen-web, m365-copilot-web, …) renders a one-click external link (opening the provider's login/home page in a new tab) so operators no longer tab away to retype the URL mid-setup. The host resolves from a pure, unit-tested resolveWebProviderHost() (prefers WEB_COOKIE_PROVIDERS[id].website, falls back to the registry baseUrl origin); non-web providers render e...
Read more

v3.8.45

Choose a tag to compare

@diegosouzapw diegosouzapw released this 06 Jul 06:03
3ddcee6

✨ New Features

  • feat(providers): add Yuanbao (web) as a cookie-session provider (#6196) — yuanbao-web (Tencent Yuanbao, yuanbao.tencent.com) with cookie-only auth (hy_user/hy_token + public agent id), SSE→OpenAI translation incl. reasoning_content, exposing DeepSeek V3/R1 + Hunyuan / Hunyuan-T1. Regression guard: tests/unit/providers-yuanbao-web.test.ts. together-web was deferred (no verifiable web-session endpoint — needs a captured request) and huggingchat-web dropped (the existing huggingchat already is a web-cookie provider). (thanks @chirag127)
  • feat(providers): route the built-in agentrouter through the dynamic Claude-Code wire image (#6056) — a small static allow-set (CC_WIRE_IMAGE_BUILTINS in open-sse/services/ccWireImageBuiltins.ts), consulted by isClaudeCodeCompatible / isClaudeCodeCompatibleProvider / applyFingerprint, makes agentrouter adopt the CC wire-image headers + fingerprint while guarding the CC baseUrl/auth branches so it keeps its own registry baseUrl and x-api-key auth. Regression guard: tests/unit/agentrouter-cc-wire-image.test.ts (asserts the wire image is applied AND agentrouter's baseUrl/auth are preserved). Live WAF-acceptance against agentrouter.org is a VPS validation follow-up (Hard Rule #18).
  • feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174) — cloudflare-ai is removed from the bulk-add exclusion list and the bulk parser gains a 3-field name|accountId|apiKey mode; the bulk route now builds a per-entry providerSpecificData so each key carries its own accountId (fixing the previous shared-object reuse), and both the create + key-validation paths receive it. Regression guard: tests/unit/bulk-api-key-parser-cloudflare.test.ts. (thanks @muflifadla38)
  • feat(dashboard): routing/settings UX clarity (#6147) — (1) weighted combos show the effective routing share % next to each weight when weights don't sum to 100 (WeightTotalBar.tsx); (2) the status widget's user-facing "Cloud Sync" label is renamed to "Remote Settings Sync" (CloudSyncStatus.tsx; internal ids/state untouched); (3) built-in providers gain an opt-in advanced base-URL override (isBaseUrlOverrideEligibleProvider, hidden behind an "Advanced" toggle, reusing the existing providerSpecificData.baseUrl persistence — not globally widened). Regression guard: tests/unit/routing-settings-ux-6147.test.ts.
  • feat(combo): add an option to disable session stickiness, per-combo or globally — round-robin / random combos can rotate to a different connection on every request instead of pinning a whole conversation to one connection by its first-message hash. Resolution precedence per-combo config.disableSessionStickiness → global settings.disableSessionStickiness → default false (preserves the #3825 prompt-cache/504 fix); gates both stickiness call sites in open-sse/services/combo.ts. Exposed as a global toggle (Combo Defaults) and a per-combo Inherit/on/off control. (#6168) Regression guard: tests/unit/combo-disable-session-stickiness.test.ts. (thanks @RCrushMe)
  • feat(docker): add the OMNIROUTE_NO_SUDO env flag for root-less / user-namespaced deployments — the MITM cert-trust command path (resolveSudoSpawn in src/mitm/systemCommands.ts) now strips the leading sudo when the flag is truthy, in addition to the existing root / sudo-missing cases, so the Proxy Agent runs without sudo (the operator trusts the CA manually, e.g. via NODE_EXTRA_CA_CERTS). Argv-array spawn preserved — no shell interpolation (Hard Rule #13). (#6122) Regression guard: tests/unit/mitm-systemCommands-no-sudo.test.ts. (thanks @powellnorma)
  • feat(providers): add Requesty as an OpenAI-compatible gateway provider (BYOK, base https://router.requesty.ai/v1, ~200 free requests/day) — wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/requesty/, src/shared/constants/providers/apikey/gateways.ts). (#6120) Regression guard: tests/unit/requesty-provider.test.ts. (thanks @chirag127)
  • feat(dashboard): add configured-only / available-only filters to the Free Provider Rankings page (#6150) — hide providers you haven't configured, or whose connections are all rate-limited / out of quota, via server-side query params (?configuredOnly / ?availableOnly on GET /api/free-provider-rankings) backed by a testable lib helper reusing the in-process connection state (no Redis). Both filters default off, so the default view is unchanged; this supersedes the earlier client-side "Configured Only" toggle (#6245) with an available-only dimension and unit-tested logic. Regression guard: tests/unit/freeProviderRankings-filters.test.ts.
  • feat(rankings): add a 'Configured Only' filter to the Free Provider Rankings page, so the table can be narrowed to just the providers you have configured connections for (with an empty-state hint when none are configured). New en.json keys and a pure filter helper covered by tests/unit/free-provider-rankings-configured-filter.test.ts. (#6245, closes #6150 — thanks @Iammilansoni)

🔧 Bug Fixes

  • fix(mitm): the test suite and CI can never mutate the OS trust store again — OMNIROUTE_SKIP_SYSTEM_TRUST=1 (set by the global test setup and all CI workflows) makes installCert/uninstallCert/installTproxyCa skip the privileged OS dispatch while preserving the #4546 environment-skip contract. Root cause of the self-hosted runner incident: a cert-flow integration test installed a 105-byte fake PEM into /usr/local/share/ca-certificates, breaking ALL system TLS on the VM. Regression guard: tests/unit/system-trust-test-guard.test.ts. (#6310)
  • fix(security): /api/keys/{id}/devices answers a clean method-first 405 for undocumented HTTP methods (e.g. the new QUERY) via a dedicated http-method-guard rule — the auth layer was answering 401 first, failing schemathesis's unsupported-methods check. Same pattern as the v3.8.44 TRACE fix. Regression guard: tests/unit/dast-method-not-allowed.test.ts.
  • fix(combo): the #6216 empty-stream failover is restricted to truly empty bodies (zero bytes — the Gemini HTTP-200-empty case), restoring the #3399/#3685 pass-through contracts for [DONE]-terminated empty streams and incomplete Claude lifecycles. New guard: #5976 truly EMPTY streaming body → invalid for combo failover (87/87 across both suites).
  • fix(combo): 5 streaming-path fixes — locked-stream 500, error-frame-only-if-no-content, Gemini MALFORMED_RESPONSE→content_filter failover, correlationId substring search, per-model-500 lockout skip + request-logger UI detail. Maintainer follow-up: releaseQualityClone cancels the abandoned quality-check tee branch (per-request memory) + regression test. (#6216 — thanks @hartmark)
  • fix(skills): generate the missing omni-github-skills registry entry (the #6186 catalog addition never ran the generator — 8 integration assertions split between old/new counts) and align the agent-skills catalog counts across integration + unit suites (43 = 23 API + 20 CLI; 44 with config).
  • fix(a2a): finish the #6186 catalog-count update — listCapabilities metadata reported coverage.api.total: 22 (type literal + value) and SkillCoverageSchema pinned z.literal(22), so the schema would REJECT the correct runtime value with 23 API skills. All three aligned to 23.
  • fix(github-skills): add a missing import, unit tests and a settings JSON-parse fix for the GitHub agent-skill discovery/import flow. (#6186 — thanks @Moseyuh333)
  • fix(api): POST /api/github-skills validates its body with a Zod schema (validateBody) instead of blind request.json() destructuring — a non-array targets would crash .map. Regression guard: tests/unit/github-skills-route-validation.test.ts.
  • fix(docker): add id= to the BuildKit cache mounts so strict builders (e.g. buildkitd with strict frontend parsing) accept the Dockerfile. (#6291 — thanks @karimalsalah)
  • fix(oauth): register zed in the OAuth PROVIDERS map (fixes "Unknown provider" on the Zed sign-in flow) (#6078 — thanks @anki1kr), and align zed in OAUTH_PROVIDER_IDS + the config enum after the merge.
  • fix(doubao-web): switch the Doubao web provider to the Dola global endpoint. (#6235 — thanks @backryun)
  • fix(doctor): resolve two false-positive WARNs in the doctor diagnostics (#6163, closes #6162 — thanks @arssnndr)
  • fix(providers): refresh the GitHub Copilot model catalog to the current upstream set. (#6154 — thanks @backryun)
  • fix(providers): correct the Kiro model catalog to real upstream ids — fabricated claude-opus-4.7/claude-sonnet-4.6 entries removed, real claude-sonnet-5/claude-sonnet-4.5/claude-haiku-4.5 kept. (#6170...
Read more

v3.8.44

Choose a tag to compare

@diegosouzapw diegosouzapw released this 04 Jul 16:26
1bda6c1

✨ New Features

  • feat(resilience): throttle upstream quota fetches on the per-request preflight path (#6009) — a new global min-interval gate (open-sse/services/quotaFetchThrottle.ts) spaces the actual network calls made by the Codex quota fetcher so that many accounts on one IP no longer fetch quota in the same second (which, per router-for-me/CLIProxyAPI#2385, can get a Codex OAuth token revoked). Complements the existing bulk-sync spacing (PROVIDER_LIMITS_SYNC_SPACING_MS) which already serialized the periodic provider-limits sync — this covers the concurrent combo/preflight path it didn't. Cache hits are never delayed; fail-open (only ever awaits a timer). Configurable via OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, clamped 0..5000; 0 disables). Regression guard: tests/unit/quota-fetch-throttle-6009.test.ts (5). (thanks @powellnorma)
  • feat(autoCombo): add per-request Auto-Combo controls via two headers (#6024 / #6025 / #6023) — X-OmniRoute-Mode steers an auto combo's scoring for a single request (friendly presets fast/balanced/quality/cheap/reliable/offline or a raw mode-pack name; balanced forces the default weights), and X-OmniRoute-Budget sets a hard per-request USD cost ceiling. Both override the combo's stored config only for the request that carries them; unknown/garbage values are ignored so the saved config is preserved. The resolvers are pure (open-sse/services/autoCombo/requestControls.ts) and feed the engine's existing config.modePack / config.budgetCap inputs — no engine changes. Regression guard: tests/unit/auto-combo-request-controls-6024.test.ts (5). (thanks @chirag127)
  • feat(providers): add the Kenari OpenAI-compatible gateway (BYOK). Regression guard: tests/unit/kenari.test.ts. (thanks @doedja)
  • feat(models): add claude-sonnet-5 to the Antigravity model catalog (alias mapping in antigravityModelAliases.ts) (#6103). Regression guard: tests/unit/antigravity-model-aliases.test.ts. (thanks @anki1kr)
  • feat(api): add /v1/ocr endpoint (Mistral OCR), an OCR provider category, and Mistral moderation support. (#5950) (thanks @waguriagentic)
  • Discovery tool (Phase 2): add the discoveryResults DB module (CRUD over the discovery_results table, migration 074) and wire the opt-in provider-discovery service to persist and read findings through it (persistDiscoveryResult, getDiscoveryResults, getDiscoveryResultById, markVerified, deleteDiscoveryResult) with (provider, method, endpoint) upsert de-duplication. Adds the /api/discovery/* HTTP surface — GET /results, GET|DELETE /results/:id, POST /scan, POST /verify/:id — under strict loopback-only authorization (/api/discovery/ is in LOCAL_ONLY_API_PREFIXES and is NOT manage-scope-bypassable, so the scan route's outbound probes can never be reached from a tunnel/remote origin). Adds a dashboard UI tab (Tools → Discovery, /dashboard/discovery) to run scans and review, verify, or delete findings. The service stays opt-in / default-off. (#5939)
  • feat(api): expose a read-only provider plugin manifest at GET /api/v1/provider-plugin-manifest for sidecar/relay discovery. (#6001) (thanks @KooshaPari)
  • feat(sidecar): advertise the provider manifest URL to Bifrost/CLIProxyAPI via the X-OmniRoute-Provider-Manifest-Url header (OMNIROUTE_PROVIDER_MANIFEST_URL). (#6007) (thanks @KooshaPari)
  • feat(autoCombo): add a latency/speed-optimized routing mode (shared rankBySpeed scoring core) plus the omniroute_pick_fastest_model MCP tool. (#6011) (thanks @KooshaPari)
  • feat(resilience): surface Codex banked reset credits per connected account (#5199) — the Codex quota parsers (buildCodexUsageQuotas, parseCodexUsageResponse) now additively read rate_limit_reset_credits.available_count (+ optional rate_limit_reached_type) from the /wham/usage payload OmniRoute already fetches, and the provider-limits dashboard renders a "Banked Reset Credits" row when a positive count is present. Display-only and fail-open — the field is eligibility-gated, so accounts without it are unaffected (parsers never throw on absent/garbage shapes); redemption (an unofficial mutating endpoint) is intentionally out of scope. Regression guard: tests/unit/codex-banked-reset-credits-5199.test.ts (8). (thanks @ofekbetzalel)
  • feat(providers): add sign-up geo-restriction notices for SenseNova and StepFun (#5462) — the provider add-form now warns that SenseNova's console appears to require a Chinese (+86) phone number with no documented international path, and that StepFun's default endpoint is its China platform while a global StepFun Open Platform (platform.stepfun.ai, operated by Sparkling AI Pte. Ltd., Singapore) with email/Google/Discord login exists for international users. Informational notice only — neither provider is disabled. Regression guard: tests/unit/regional-provider-cn-notices-5462.test.ts. (thanks @chirag127)
  • feat(usage): add on-demand period-scoped usage-data reset (Settings → System Storage) with a purge API and time-window selector. (#5831)
  • feat(claude-code): add an opt-in auto-permission classifier compat mode (off/auto/always) for Claude Code, toggleable from the CLI Code settings. (#5810)
  • feat(providers): add optional client-identity header profiles for compatible nodes — preset User-Agent/fingerprint headers (e.g. matching a known CLI) merged into the existing customHeaders field. (#5812)
  • feat(build): add a backend-only fast build mode (scripts/build/build-next-isolated.mjs + backendOnlyPages.mjs) that skips compiling the dashboard frontend pages, cutting local/CI build time for backend-only changes. (#6119 — thanks @artickc)
  • feat(minimax): extract MiniMax M3's raw <think>...</think> leakage into reasoning_content on the 8 OpenAI-format provider tiers, leaving the Claude-format minimax/minimax-cn tiers untouched (they already report reasoning correctly). (#6073 — thanks @KooshaPari)
  • feat(services): promote Bifrost (@maximhq/bifrost — Go AI-gateway) from an env-only relay sidecar to a first-class embedded/supervised service, matching the existing cliproxy/9router model — installer, bootstrap SERVICES[] entry, migration 113 DB seed, 7 lifecycle API routes under /api/services/bifrost/ (loopback-only), a dashboard tab, and relay auto-wiring that defaults BIFROST_BASE_URL to the supervised port when running. Implements item #2 of #5670; the broader RouterBackend contract (items #1, #3-#5) stays out of scope. (#5817, part of #5670)
  • feat(services): add Mux (coder/mux — local agent-orchestration daemon) as a fourth-tier embedded service on the existing ServiceSupervisor framework — npm-based installer, bootstrap.ts registration, migration 113 DB seed, 7 lifecycle API routes under /api/services/mux/ (loopback-only, defense-in-depth bind to 127.0.0.1), and a dashboard tab reusing the shared service-management components. (#6034)
  • feat(xai): surface Grok/xAI usage on the quota dashboard via local usageHistory aggregation (getXaiUsage) — since xAI exposes no per-account quota API, this sums tokens routed to the connection from usage_history and reports them as a cumulative, uncapped quota, mirroring the existing Xiaomi MiMo self-track pattern. (#5806)
  • feat(minimax): extract MiniMax M3's raw <think>...</think> tags into a separate reasoning_content field on the 8 provider tiers that register M3 with format:"openai" (trae, huggingchat, bazaarlink, ollama-cloud, opencode, cline, opencode-zen, codebuddy-cn) — previously the thinking text leaked directly into content. Reuses the existing extractThinkingFromContent primitive, extending its allowlist with a minimax-m3-only pattern; the two direct minimax/minimax-cn tiers are untouched since they already surface reasoning natively over Anthropic's Messages format. (Inspired by 9router#2231.) (#6050 — thanks @KooshaPari)
  • feat(i18n): auto-detect the browser language on first visit — a pure detectBrowserLocale() matcher (exact match, zh-HK/zh-MO folded to zh-TW, language-prefix match, else null) plus a client-only LocaleAutoDetect component mounted once in the root layout. When no locale cookie is set yet, it reads navigator.languages, computes a match against the supported locales, and persists it via the same cookie/localStorage writer LanguageSelector already used (extracted to shared/lib/persistLocale.ts). (Inspired by 9router#1324.) (#5979)
  • feat(cli-tools): add CodeWhale — the actively-maintained successor to DeepSeek TUI (same author, renamed pr...
Read more

v3.8.43

Choose a tag to compare

@diegosouzapw diegosouzapw released this 02 Jul 15:43
b729a8f

[3.8.43] — 2026-07-02

✨ New Features

  • usage (quota percentages + provider USD drilldown): @@om-usage and the HTTP usage endpoint now report personal API-key quotas as remaining percentages (USD amounts stay out of the command output), provider quota remaining is scaled by the configured quota cutoff so the protected reserve reads as 0% left, and the quota dashboard regains a provider USD cost drilldown (/api/usage/provider-window-costs + ProviderUsdCostModal, management-auth gated). Also honors observed provider quota resets: a same-resetAt reset (usage dropping back to the reset floor) is detected and preferred over stale recorded weekly events for provider USD windows and API-key USD quotas. New src/lib/usage/providerWindowCosts.ts. Regression guards: tests/unit/provider-window-costs.test.ts, tests/unit/internal-usage-command.test.ts, tests/unit/api-key-usage-limits.test.ts, tests/unit/lib/quota-reset-events.test.ts. Extracted from #5863 by @Witroch4.

  • dashboard (live WS behind reverse proxy): the live dashboard WebSocket can now be fronted by a reverse proxy or Cloudflare Tunnel via NEXT_PUBLIC_LIVE_WS_PUBLIC_URL (e.g. wss://ws.my-ai.com/live-ws). The URL is honored both at build time (env inlined into the bundle) and at runtime for prebuilt Docker/npm images: the /api/v1/ws?handshake=1 handshake now echoes a lazily-read live.publicUrl (only ws:///wss:// values are accepted; anything else is rejected to null), and useLiveDashboard resolves the URL from that handshake before connecting, falling back to the previous ws(s)://hostname:20129 default. Also documents LIVE_WS_ALLOWED_HOSTS and aligns the GitLab Duo OAuth scopes line in .env.example with the live config (ai_features read_user). Regression guard: tests/unit/live-ws-public-url.test.ts (5). (#5877 by @ianriizky)

  • providers (CLI profile auto-sync): opt-in toggles to auto-regenerate CLI tool profiles after a provider model sync. When enabled, a model-catalog change (re)writes that tool's profile files from the live catalog — Codex (~/.codex/*.config.toml) and now Claude Code (~/.claude/profiles/<name>/settings.json, via an extracted syncClaudeProfilesFromModels + a new claudeProfileAutoSync.ts mirroring the Codex path). Both are off by default and never touch the active/default CLI config; they are backed by the OMNIROUTE_AUTO_SYNC_CODEX_PROFILES / OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILES feature flags (DB/dashboard override > env > default "false") and additionally gated behind the existing CLI_ALLOW_CONFIG_WRITES write-guard. A "CLI profile auto-sync" card on the CLI Code dashboard toggles each (moved from the providers dashboard in #5778 — thanks @rdself). Regression guards: tests/unit/claude-profile-auto-sync-gate.test.ts, tests/unit/codex-profile-auto-sync-gate.test.ts, tests/unit/cli/setup-claude.test.ts (follow-up to #5737).

  • cli (startup banner): the serve startup banner now prints the running OmniRoute version (v3.8.x) beneath the ASCII logo, so the active version is visible at a glance without a separate --version call. Regression guard: tests/unit/cli-serve-version-banner.test.ts. Thanks @chirag127 (#5752).

  • analytics (subscription cost): flat-rate providers now show $0 in cost analytics instead of an inflated per-token estimate. Subscription / coding-plan providers (every cookie-web provider — ChatGPT Web, grok-web, … — plus the dedicated Minimax Coding, Kimi Coding, GLM Coding, Alibaba Coding Plan, and Xiaomi MiMo plans) bill a flat fee, not per token, yet still carry per-token pricing rows used for estimates — so the analytics dashboard over-reported their cost. A new flat-rate classifier (src/lib/usage/flatRateProviders.ts) is consulted by the analytics surfaces (analytics route, usage stats, usage analytics) via an opt-in flatRateAsZero cost option, so those providers read $0 while budget / quota / routing keep estimating unchanged. Deliberately NOT zeroed: codex/cx (OmniRoute actively tracks Codex token cost — Fast-tier multipliers, GPT-5.x pricing — and Codex can be a metered account), byteplus (metered ModelArk), minimax-cn (metered China API). Regression guard: tests/unit/flat-rate-cost-5552.test.ts. (#5552)

  • mcp (RTK): expose the RTK tool-output learn/discover workflow as two new MCP tools so an agent can grow the RTK filter catalog without leaving the protocol. omniroute_rtk_discover analyzes recently captured raw tool output (discoverRepeatedNoise / suggestFilter) and returns candidate noise patterns plus a suggested filter; omniroute_rtk_learn lists the captured command samples (listRtkCommandSamples) and resolves a command to its RTK filter id (commandToId). Both are read-only (scope read:compression), wrap the existing RTK discovery primitives (no new logic in the engine), and log to the MCP audit trail. Regression guard: tests/unit/compression/rtk-mcp-tools.test.ts (4). gaps v3.8.42 — T07.

  • compression (LLM tier): add an opt-in, default-off LLM-tier compression engine (llm) that condenses the prose of non-system messages via a pluggable chat-completion backend. It mirrors the llmlingua engine's contract but is safe by construction: the default backend is a no-op pass-through (the engine never mutates the payload until an operator both enables it and wires a real backend via setLlmCompressorBackend()), it is not part of the default stacked pipeline, enabled defaults to false, fenced code blocks and system messages are never sent to the model, and every backend error fails open (the original segment/body is kept, never thrown). A minTokens floor skips small prompts. The real production backend is intentionally a VPS-validated follow-up (Hard Rule #18), exactly as the llmlingua worker backend is gated. New open-sse/services/compression/engines/llm/index.ts. Regression guard: tests/unit/compression/llm-compressor-engine.test.ts (8). gaps v3.8.42 — T05/C3.

  • memory (typed decay): add opt-in typed memory decay (TV6) so the conversational memory store stops accumulating stale episodic noise. Each injected memory now tracks an access_count + last_accessed_at (always-on, non-destructive telemetry; migration 111_memory_typed_decay), and an opt-in, default-off sweep (MEMORY_TYPED_DECAY_ENABLED, default false) deletes memories that are past a per-type TTL and not immune. Only episodic decays by default (30d, env-tunable); factual/procedural/semantic are immune, and any memory accessed >= 3 times earns access immunity (mirroring "guardrail/convention/decision never decay"). The decay clock re-bases on the last access, so used memories survive. Deletions reuse deleteMemory (SQLite + sqlite-vec + Qdrant stay in sync) and fail open; an optional periodic sweep is doubly opt-in (also needs MEMORY_TYPED_DECAY_SWEEP_INTERVAL>0). With the flag off nothing is ever deleted (Rule #20 spirit). New src/lib/memory/typedDecay.ts. Regression guard: tests/unit/memory/typed-decay.test.ts (15). gaps v3.8.42 — T10/TV6.

  • dashboard (combos): the named-combos editor now lets you drag to reorder the stacked-compression pipeline instead of only editing fixed-position steps. A new pure model (src/shared/components/compression/compressionPipelineModel.ts) owns add/remove/move/update with the engine→intensity invariant and a never-empty guarantee, and a @dnd-kit/sortable editor (CompressionPipelineEditor.tsx, matching the sidebar reorder pattern) replaces the inline list in CompressionCombosPageClient. Order persists through the existing combos endpoint. Regression guards: tests/unit/compression-pipeline-model.test.ts (11) + tests/unit/ui/compression-pipeline-editor.test.tsx (4). A dedicated tests/e2e/compression-studio.spec.ts (Tela A render + tab switch) closes the studios e2e gap the combo-live spec did not cover. gaps v3.8.42 — T06 + T03.

  • compression (pipeline): add an opt-in, default-off per-engine circuit-breaker to the stacked compression pipeline (T02). When an engine throws repeatedly across requests, its breaker opens and the stacked loops skip that engine (keeping the body verbatim for that step — fail-open) for a cooldown, then probe once (lazy half-open); success closes it, a failed probe re-opens it. This is distinct from the provider circuit-breaker (src/shared/utils/circuitBreaker.ts, provider-scoped + DB-persisted) — the new pipelineEngineBreaker.ts is engine-scoped, process-local, and adds zero DB/IO on the hot path. It composes with the existing per-request TV1 bail-out (which skips within a single request); the breaker adds cross-request memory. Default off (COMPRESSION_PIPELINE_BREAKER_ENABLED=false) → byte-identical to the pre-breaker pipeline (a throwing engine still propagates unless TV1 is separately enabled). Configurable per-call, per-CompressionConfig, or via env (_THRESHOLD/_COOLDOWN_MS). Regression guard: tests/unit/compression/pipeline-circuit-breaker.test.ts (9, incl. a throwing-engine integration); existing strategySelector/bail-out suites stay green. gaps v3.8.42 — T02 (2.2).

  • compression (CCR): the CCR retrieval-feedback (H8) is now graduated instead of a binary cliff. Previously a block retrieved >= 3 times was flagged do-not-compress and everything below that stayed fully compressible. Now each prior retrieval raises a block's effective minChars linearly (`effectiveMinCh...

Read more