[Improve] Gate the linked-issue fetch on the routing precheck; default routing to Gemini 3.6 Flash#716
Merged
Merged
Conversation
…ault routing to Gemini 3.6 Flash The external issue fetch added in #693 ran before every routing call with a pasted GitHub/Linear issue link, paying up to the 8s lookup deadline even when the message alone already routed. Routing now runs a tool-free precheck first and only fetches when it returns needsExternalLookup=true, then re-routes with the issue as untrusted context. Fail-open behavior is unchanged: an unavailable integration or empty fetch keeps the precheck decision. Routing inference now resolves context.routingModel, then the new R_ROUTER_MODEL deployment override, then defaults to google/gemini-3.6-flash instead of falling through to the deployment small model. In the router external-lookup gate eval (4 repetitions, 5 candidates), Gemini 3.6 Flash was the only stable gate: it asked for the issue exactly when the message alone couldn't route (0% false lookups, 100% recall) at ~4.5s per call. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
|
1 issue outstanding. See task
Reviewed 65087fd |
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Follow-up to #693.
Two-step gate.
gatherContextFromConfiguredMcpsnow runs a tool-free routing precheck first and only callsgatherExternalIssueContextwhen the precheck returnsneedsExternalLookup=true— finally wiring the signal that has been telemetry-only since the old two-step design. When context is fetched, routing runs once more with the issue as untrusted reference material and reportsphase: 'mcp'with the tools used. Fail-open is unchanged: an unavailable integration, inaccessible issue, or empty fetch keeps the precheck decision (phase: 'direct').Routing model default. Workspace routing (Slack/Linear/Discord/etc. and GitHub paths) now resolves its model as
context.routingModel→ newR_ROUTER_MODELdeployment override →google/gemini-3.6-flash, instead of falling through to the deployment-wide small model. The decision's debugmodelfield now reports the actual resolved model instead of theroomote-small-modelplaceholder.Why
#693 paid the issue-fetch deadline (up to 8s) on every routing call containing a pasted GitHub/Linear issue link, even when the message alone already routed. In the router external-lookup gate eval (24 labeled cases × 4 repetitions × 5 candidate models, run against the production routing prompt), roughly two-thirds of link-bearing messages routed correctly with no fetch — and Gemini 3.6 Flash was the only stable gate: 0% false lookups and 100% lookup recall in every repetition, at ~4.5s p50 per call. Fetched issue context took genuinely ambiguous links from ~50% blind accuracy to 100% for every model, so the fetch stays — it just only runs when it can change the decision.
Trade-off: when a lookup is needed, routing now costs two model calls plus the fetch (one extra call vs #693). The eval says that's the minority of link-bearing tasks.
Tests
phase: 'mcp').phase: 'direct'.vitest run src/server/router: 12 files, 94 tests green; slack router-debug/auto-route consumers green;tsc --noEmitclean.🤖 Generated with Claude Code