v1.2.0 #2809
missBerg
announced in
Announcements
v1.2.0
#2809
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Agent Router v1.2.0
Agent Router v1.2.0 is the first release under the project's new name (formerly Envoy AI Gateway). Nothing you deploy is renamed: the aigateway.envoyproxy.io API group, the CRDs, the Helm charts, the container images, and the aigw CLI all stay the same. This release hardens MCP authentication and cross-namespace references. It adds bare tool names, caller-supplied credentials, and policy merging to MCPRoute, plus two new backend schemas (AWSOpenAI and TypeSafe). It also moves to Envoy Gateway v1.9.2, Gateway API v1.6, and Gateway API Inference Extension v1.6. Several security fixes change behavior, so read the breaking changes before you upgrade.
✨ New Features
Agent Router
MCP Gateway
prefixMode: Never— Clients that hardcode tool names, such as interactive MCP Apps, no longer have to see the<backend>__prefix. SetprefixMode: Neveron the route or on individualbackendRefs. A Never-mode backend must list its tools intoolSelector.include, and the controller marks the route NotAccepted if two Never-mode backends expose the same name. Prompts are exposed bare only when they are listed in the newpromptSelector.include. Resources always keep the prefix. The default is stillAlways.injectionPolicy: IfNotPresent— Let users bring their own upstream token, for example a personal GitHub PAT forwarded withforwardHeaders, and fall back to a shared service key when they don't. WithIfNotPresent, the configured API key is injected only when the target header is missing. This applies to header injection only.mergeType— The SecurityPolicy and BackendTrafficPolicy that Agent Router generates for an MCPRoute used to replace any Gateway- or listener-level policy. SetsecurityPolicy.mergeTypeorbackendTrafficPolicy.mergeTypetoStrategicMergeorJSONMergeto keep a Gateway-wide rate limit or ext-auth policy in force alongside MCP OAuth.oauth.authorizationServerMetadataUrlpoints the controller at the exact RFC 8414 metadata document when the issuer URL doesn't lead to it. The controller fetches this document while reconciling, so it needs network access to the identity provider. If the fetch fails, reconciliation fails.initializeresponse now negotiates the protocol version with the client and backends instead of always answering 2025-06-18. Merged list responses also carryttlMsandcacheScopecaching hints. Support for the stateless 2026-07-28 MCP specification is in progress and not yet enabled.Providers & API Compatibility
AWSOpenAI) — Send OpenAI-format/v1/chat/completionsand/v1/responsesrequests to Bedrock Runtime's OpenAI-compatible/openai/v1endpoint without a translation hop. Requests are signed with SigV4 through an AWSBackendSecurityPolicy. The prefix defaults toopenai/v1. Other endpoints return 422.TypeSafe) — Route TypeSafe's Jev decision model through the gateway at/typesafe/v1/systemonewith API-key auth. Request bodies pass through unchanged apart from a model name override, and token usage feeds the standard metrics andllmRequestCosts. Streaming is not supported.anthropic-betafiltering — Oneanthropic-betavalue that a provider rejects no longer fails the whole request.AIServiceBackend.spec.headerValueFiltersdrops listed values (Denylist, the default) or keeps only listed values (Allowlist). It applies to/v1/messagestraffic toGCPAnthropicandAWSAnthropicbackends. Bedrock also accepts four more beta flags, includingthinking-token-count-2026-05-13.reasoning_efforton Vertex AI and Bedrock, and structured output (response_format) on Vertex AI. Chat-completion tools forwardeager_input_streamingto Claude. Gemini on Vertex AI accepts vLLM-stylestructured_outputs(json,regex,choice).cache_write_tokens— Prompt-cache writes reported ascache_write_tokensby OpenAI Chat Completions and the Responses API now count toward cache-creation token metrics and costs. Translated Claude and Bedrock responses report the same value.Model Catalog & Observability
/v1/models— SetexcludeFromModelsEndpoint: trueon anAIGatewayRouterule to keep internal or backwards-compatible model aliases out of/v1/modelsand/anthropic/v1/models. Requests to those models still route normally.gen_ai.*metrics now carry agen_ai.backendattribute (gen_ai_backendin Prometheus) set to the backend'snamespace/name. This lets you compare latency, time to first token, and token usage across providers serving the same model.Inference & Operations
InferencePoolwithoutendpointPickerRefis now reported asAccepted=False(EndpointPickerRefMissing) instead of breaking routing.appProtocol: kubernetes.io/h2cis honored for cleartext HTTP/2 model servers.controller.priorityClassNamevalue.🔗 API Updates
AIGatewayRouteRule.excludeFromModelsEndpoint— Optional boolean (defaultfalse) in v1alpha1 and v1beta1.AIServiceBackend.spec.headerValueFilters— Optional list (max 16, one per header) of{name, mode, values}.modeisDenylist(default) orAllowlist;valuesholds up to 64 entries. A CEL rule rejects the field unless the schema isGCPAnthropicorAWSAnthropic, and onlyanthropic-betais honored today. v1beta1 only.VersionedAPISchema.name:AWSOpenAIandTypeSafe— Two new schema values.TypeSafedefaults its version tov1.GatewayConfig.spec.extProc.metadataForwardingNamespaces— Optional list (max 32) of dynamic metadata namespaces that Envoy forwards to the external processor. It is required forcredentialOverride.fromDynamicMetadata. Under Envoy GatewaymergeGateways, every Gateway in the class shares the union of these namespaces.MCPRoute.spec.prefixModeandbackendRefs[].prefixMode—AlwaysorNever; unset meansAlways. The per-backend value takes precedence. v1beta1 only.MCPRoute.spec.backendRefs[].promptSelector— Filters prompts withinclude,includeRegex,exclude, orexcludeRegex(max 32 each). Exclude rules win over include rules. v1beta1 only.MCPRoute.spec.securityPolicy.mergeTypeandspec.backendTrafficPolicy.mergeType— Envoy GatewayMergeType;Replaceis rejected. Unset keeps the previous override behavior.MCPRoute.spec.backendRefs[].securityPolicy.apiKey.injectionPolicy—Always(default) orIfNotPresent.IfNotPresentcannot be combined withqueryParam. v1beta1 only.MCPRoute.spec.securityPolicy.oauth.authorizationServerMetadataUrl— Optional URI, up to 1024 characters.MCPRoutevalidation for JWT-based CEL —securityPolicy.oauthis now required when an authorization rule orbackendSelectorrule'scelexpression referencesauth.jwt. See Breaking Changes.Deprecations
cache_creation_input_tokensin translated OpenAI-format responses — Responses translated from Claude and Bedrock now reportcache_write_tokensalongside the legacycache_creation_input_tokensfield. Readcache_write_tokensinstead; the legacy field will be removed in v1.3.0.default-insecure-seed. It generates a random seed, stores it in the Secret<controller fullname>-mcp-session-encryption, and reuses it on later upgrades. Upgrading therefore ends active MCP sessions, and clients must re-initialize. To keep sessions during the upgrade, setcontroller.mcp.sessionEncryption.fallback.seed=default-insecure-seedand remove it once clients have reconnected. If you render the chart withhelm template(common with GitOps), setcontroller.mcp.sessionEncryption.existingSecretorseed. Otherwise every render generates a new seed.oauth.issuer— The generated JWT provider now checks the token'sissclaim. Tokens with a missing or different issuer, including a trailing-slash mismatch, get 401. Make suresecurityPolicy.oauth.issuermatches your identity provider'sissexactly.securityPolicy.oauth— The gateway readsrequest.auth.jwtclaims and scopes, and the session subject, only from a JWT that Envoy has verified. WithoutsecurityPolicy.oauth, claims are empty, and new CRD validation rejects authorization orbackendSelectorCEL that referencesauth.jwt. Existing routes that break this rule keep their stored spec, but every update is rejected. At runtime their claim lookups fail, and therefore deny, and their scope checks evaluate to false. Addoauthor remove the JWT references. MCP sessions are now bound to the verified subject; reusing a session ID as a different user returns 401.backendSelectorCEL expression that fails at runtime, for example by reading a header that isn't present, used to be skipped. It now denies:tools/callreturns 403, the tool is hidden fromtools/list, and the backend is left out of the session. Write expressions defensively, for example"x-team" in request.headers && request.headers["x-team"] == "blocked".spec.headers— On an MCPRoute with bothheadersmatches andoauth, the OAuth well-known endpoints now apply the same header matches. Clients that run OAuth discovery without those headers get 404.BackendSecurityPolicyneed a ReferenceGrant —azureCredentials.clientSecretRef,gcpCredentials.credentialsFile.secretRef, and the OIDCclientSecretunder AWS, Azure, and GCPoidcExchangeTokenwere read from other namespaces without any check. ThesecretRefofapiKey,azureAPIKey,anthropicAPIKey, andawsCredentials.credentialsFileignored its namespace, so a cross-namespace reference failed or silently used a Secret with the same name from the policy's own namespace. All of these now honor the namespace, and a missing grant sets the policy to NotAccepted. Already-issued tokens keep working until they expire, so the failure may appear late. Create a ReferenceGrant in the Secret's namespace (see Upgrade Guidance).AIGatewayRoutebackendRefto anAIServiceBackendorInferencePoolin another namespace without a ReferenceGrant already marked the route NotAccepted, but its credentials were still written to the external processor's config. Such backends are now left out. Working setups already have the grant; this affects only routes that were already NotAccepted.to.nameis now enforced — A ReferenceGranttoentry that setsnameused to authorize every object of that group and kind in its namespace. It now authorizes only the named object. References to other objects, such as another Secret,AIServiceBackend, orInferencePool, are NotAccepted and left out of the data plane. Check grants that setnamebefore upgrading.QuotaPolicyDistinctfix, tenants can start receiving 429 at the limits you configured. Review per-tenant limits before upgrading.-configPathremoved — The external processor reads only its config bundle (-configBundlePath), and the controller no longer writes the legacyfilter-config.yamlSecret. Existing legacy Secrets are left in place and can be deleted. This matters only if you run the external processor yourself, or still have Envoy pods created before config bundles existed (v1.0.x era). Restart those pods before upgrading.🐛 Bug Fixes
BackendSecurityPolicythat targets anInferencePoolnow also reach the Gateway's filter config.fromentry, now reconciles the affectedAIGatewayRoutes (includingInferencePoolbackends) andBackendSecurityPolicys (including OIDC client secrets). Access they relied on used to stay in place until an unrelated reconcile.backendRefbefore itsAIServiceBackendexisted could panic and crash-loop the controller. The controller now logs the mismatch and continues.helm upgrade— With the default self-signed webhook certificate, an upgrade or Argo CD sync wiped the webhook'scaBundle, so Envoy pods failed admission until the controller restarted. The chart now embeds the CA bundle.credentialOverride.fromDynamicMetadatanow receives metadata — In v1.1.0 Envoy never forwarded the metadata namespace to the external processor. Requests silently used the configured credential, or failed with 401 whenfallbackToConfigured: false. List the namespace inGatewayConfig.spec.extProc.metadataForwardingNamespaces(see Upgrade Guidance).stream_optionsin a streaming request could stopinclude_usagefrom reaching the provider, which bypassed token accounting. Usage metadata is now also recorded when a client disconnects right after the final SSE frame. Responses API streams framed with CRLF now report usage, model, and tracing data./v1/responsesto Vertex AI Gemini) returns 422 Unprocessable Entity. Oversized JSON schemas sent to Vertex AI return 422 once they exceed 50,000 nodes.QuotaPolicyDistinctheader selectors, the end-of-stream token cost went into one shared bucket instead of the caller's own. See Breaking Changes for the effect on tenants.web_search_20260209tool on/v1/chat/completionsis now forwarded to Claude instead of rejected with 422; use streaming, because non-streaming responses currently return only the first text block. Tracing of long Claude streams no longer allocates quadratically or panics on malformed delta indices.application/json; charset=utf-8(for example the MCP Java SDK) no longer have their list results dropped. OAuth metadata discovery now treats an empty or non-JSON document as a miss and tries the next well-known URL.aigw runbase URL handling — A path inOPENAI_BASE_URLorANTHROPIC_BASE_URL(for example OpenRouter's/api/v1) is now used as the request prefix instead of being dropped. Base URLs that aren't http or https are rejected.InferencePoolroutes are rejected — When Envoy Gateway couldn't resolve anInferencePool, the extension server used to accept a route that only returned 500. It now rejects that route so the error is visible.📖 Upgrade Guidance
Upgrading from v1.1 needs no CRD storage migration, and every new field is optional. Several security fixes change behavior, though, so work through the steps that apply to you before you upgrade. Upgrade from v1.1.x, not directly from v1.0.x.
1. Upgrade Gateway API and Envoy Gateway first
Envoy Gateway v1.9 requires the Gateway API v1.6 CRDs. If you use the standard channel, move any
TCPRouteandUDPRoutemanifests togateway.networking.k8s.io/v1before you upgrade the CRDs. Then upgrade the CRDs and Envoy Gateway. The Agent Router values file is unchanged:helm template eg-crds oci://docker.io/envoyproxy/gateway-crds-helm \ --version v1.9.2 \ --set crds.gatewayAPI.enabled=true \ --set crds.envoyGateway.enabled=true \ | kubectl apply --force-conflicts --server-side -f - helm upgrade -i eg oci://docker.io/envoyproxy/gateway-helm \ --version v1.9.2 \ --namespace envoy-gateway-system \ -f https://raw.githubusercontent.com/theagentrouter/agent-router/v1.2.0/manifests/envoy-gateway-values.yamlEnvoy Gateway v1.9 has its own breaking changes. Read its v1.9.0 and v1.9.2 release notes. The ones most likely to affect AI traffic:
EnvoyExtensionPolicyis disabled by default. Enable it withextensionApis.enableLua.mergeTypeon aSecurityPolicyorBackendTrafficPolicyis only accepted on route targets.Do not enable Envoy Gateway's new
EnvoyProxymergeBackendsoption with Agent Router yet. It shares clusters between routes, which breaks how Agent Router resolves the backend for a request.2. Plan for the MCP session seed change
The chart now generates a random MCP session encryption seed, so upgrading ends active MCP sessions. Pick one:
Accept a reconnect. Do nothing. Clients re-initialize their sessions.
Keep sessions during the upgrade. Add the old seed as a fallback in your values for this upgrade, then remove it once clients have reconnected:
GitOps or
helm template. The chart can't read back a seed it generated earlier, so it generates a new one on every render. Supply your own Secret instead:kubectl create secret generic mcp-session-seed -n envoy-ai-gateway-system \ --from-literal=seed="$(openssl rand -hex 32)"Then set
controller.mcp.sessionEncryption.existingSecret: mcp-session-seed.3. Check your MCPRoutes
securityPolicy.oauth.issuermust match your identity provider'sissclaim exactly, including any trailing slash.backendSelectorCEL that referencesrequest.auth.jwtneedssecurityPolicy.oauthon the same route."x-team" in request.headers && request.headers["x-team"] == "blocked"instead ofrequest.headers["x-team"] == "blocked".headersandoauth, make sure clients send those headers during OAuth discovery.4. Add ReferenceGrants for cross-namespace credential Secrets
If a
BackendSecurityPolicyreferences a credential Secret in another namespace, create a grant in the Secret's namespace. This now applies to every credential type: API keys, AWS credentials files, Azure client secrets, GCP credentials files, and OIDC client secrets.A
toentry withnamenow authorizes only that object. Withoutname, it covers every object of that kind in the namespace. If an existing grant setsname, make sure it lists every Secret,AIServiceBackend, orInferencePoolthat other namespaces reference.5. Turn on metadata forwarding for
credentialOverride.fromDynamicMetadataIf you use per-request credentials from dynamic metadata, list the metadata namespace in a
GatewayConfig, and reference thatGatewayConfigfrom the Gateway with theaigateway.envoyproxy.io/gateway-configannotation:6. Upgrade Agent Router
The chart and image names keep the
ai-gatewayprefix. Only the project name changed.7. After the upgrade
gen_ai_backendlabel. Panels that don't aggregate withsum by (...)will split into one series per backend.filter-config.yamlSecrets. You can delete them.InferencePool, upgrade the Gateway API Inference Extension CRDs to v1.6.x and redeploy your endpoint picker.📦 Dependency Versions
🙏 Acknowledgements
v1.2 came from 39 contributors, many of them contributing for the first time. Special thanks to:
🔮 What's Next
Work already in flight for upcoming releases:
MCPBackendCRD, decoupling backend configuration fromMCPRoute.The roadmap is community-driven. Join us and help shape it.
This discussion was created from the release v1.2.0.
All reactions