Repository navigation
Dynamo v1.5.1 - Release Notes
Summary
Dynamo v1.5.1 is a patch release on top of v1.5.0. It fixes Dynamo Router overload recovery, min_tokens on tokenizer-free SGLang decode workers, Qwen3-VL video routing across mixed workers, and reinforcement learning worker discovery on Kubernetes. It bounds client-supplied multimodal input (inline data: URL size, remote media downloads, image dimensions, and base64 audio) across the SGLang, TensorRT-LLM, vLLM, and vLLM-Omni paths, and moves all media fetching onto a single aiohttp client. It also upgrades the frontend runtime libraries to fix Anthropic Messages API requests from Claude Code against strict chat templates, moves ModelExpress to v0.6.0, pins PyNvVideoCodec at 2.2.3, and moves the in-tree FFmpeg build to 9.0.1.
Two changes need action on upgrade. See Breaking Changes.
Base Branch: release/1.5.1
Breaking Changes
- Egress Proxy Opt-In: Workers that fetch request media through an ambient HTTP proxy must now set
DYN_MM_TRUST_EGRESS_PROXY=1(#14474). Without it, a policy-checked media fetch that would route through a proxy is refused, because the proxy resolves the destination outside Dynamo's address check. Fetches that go direct, including hosts listed inNO_PROXY, are unaffected, and the check does not run underDYN_MM_ALLOW_INTERNAL=1. - SGLang Diffusion Local References: Local image and video files passed as
input_referenceto SGLang image-diffusion and video-generation workers now requireDYN_MM_LOCAL_PATHto name the allowed directory (#14435). Previously any local path was accepted.
Bug Fixes
- Router Overload Hint Expiry: Fixed a Dynamo Router condition where a worker that rejected requests for lack of capacity stayed marked overloaded after the pressure ended, so clients kept receiving HTTP 529 from an idle worker (#14210). Overload hints raised on the request path now expire after one second, while overload state reported by the worker monitor stays independent, so a drained worker returns to rotation without a restart.
- SGLang Tokenizer-Free Min Tokens: Fixed requests with a positive
min_tokensfailing on SGLang decode workers started withskip_tokenizer_init, where SGLang rejects themin_new_tokensfield (#14276). Dynamo now enforces the minimum itself on that path, and cancellation or early stream closure aborts every unfinished choice in multi-choice (n) requests. Requests withoutmin_tokensand workers that keep a tokenizer are unchanged. - Split Stop-Sequence Leak: Fixed hidden stop sequences that arrive split across several decode steps leaking their leading fragments into client output (#14378). The decoder now holds back text that could still complete a stop sequence and releases it only once it cannot match, or when the engine finishes on its own, for example at
max_tokens. - Qwen3-VL Video Contract Agreement: Fixed a Qwen3-VL deployment serving video with a prompt-expansion contract that only some of its workers publish (#14624). Replicas with different installed packages or engine flags can publish different video contracts; when the workers in a group disagree, Dynamo now disables exact video routing for that group while text serving continues, and keeps the existing group serving until a rebuilt one commits.
- Sidecar Worker Namespace Suffix: Fixed vLLM sidecar and mocker workers registering under the base namespace while reinforcement learning discovery on Kubernetes searched the suffixed one, so RL workloads could not find their workers (#14955). Worker registration now applies
DYN_NAMESPACE_WORKER_SUFFIXat most once, which keeps operator-upgraded workloads whoseDYN_NAMESPACEalready carries the suffix working, and an explicit command-line namespace still takes precedence. - GPU Memory Service Admission Timeout: Fixed the vLLM GPU Memory Service worker ignoring a configured
gms_ro_connect_timeout_msduring its initial weights load, which could leave startup waiting indefinitely behind lock contention (#14877). Deployments that do not set the timeout keep the existing unbounded wait. - Inline Data URL Size Cap: Added a size cap on inline
data:media URLs, enforced by both the frontend and the workers (#14437). Adata:URL larger thanDYN_MM_MAX_DATA_URL_MB(default 16 MiB) now returns HTTP 400 with a message naming the variable, where before any size was accepted. To raise the limit, set the variable on both the frontend and the workers. - SGLang Diffusion Input References: Fixed SGLang image-diffusion and video-generation workers handing a client-supplied
input_referenceto the generator after only a non-empty check (#14435). Dynamo now validates the reference and downloads a remote one to a temporary local file first, capped byDYN_MM_MAX_FILE_SIZE_MB(default 64 MiB), which matches the vLLM-Omni and TensorRT-LLM workers. - vLLM-Omni Input Validation: Fixed non-numeric or out-of-range image
widthandheightvalues and malformed base64 audio crashing the vLLM-Omni worker or ending a realtime session (#14436). Image dimensions outside 1 to 4096 are now rejected with a clear error, an unusable realtime audio chunk emits aninvalid_audioerror event instead of closing the connection, and percent-encodeddata:reference audio for text-to-speech now decodes correctly instead of producing corrupted audio. - Media Fetch Address Pinning: Fixed media fetches resolving a hostname once to validate it and again to connect, so a server that answered differently the second time was checked on one address and dialed on another (#14474). The Python workers and the Rust frontend now connect only to the addresses that passed validation, while TLS still verifies the original hostname.
- Single aiohttp HTTP Backend: Removed the opt-in httpx backend from Dynamo's HTTP client, leaving aiohttp as the only backend (#14563). Setting
DYN_HTTP_BACKEND=httpxnow logs a warning and uses aiohttp. The TensorRT-LLM multimodal processor's remote.safetensorsdownload now streams asynchronously, so it no longer blocks the event loop for up to the 300-second download timeout.
Dependency Changes
- Frontend Runtime Libraries: Upgraded
dynamo-rendererto v5.1.2 anddynamo-tokenizersto v1.8.1, fixing Anthropic Messages API failures when Claude Code sends a non-leading system message to a model whose chat template requires system content first (#14595). - DeepSeek V4 Reasoning Effort: The renderer now defaults an omitted or invalid reasoning effort to
high, treatsnoneas thinking disabled, and lets a top-levelreasoning_effortoverride the template argument (#14595). - OpenAI Response Fields: Upgraded
dynamo-protocolsto v5.4.1, so/v1/chat/completionsresponses omit absent optional fields such asusageinstead of returningnull(#14154). Client code that indexes raw response dictionaries should check for those keys. - ModelExpress: Upgraded ModelExpress to v0.6.0 in the SGLang and vLLM runtime images (#15026).
- PyNvVideoCodec and FFmpeg: Pinned PyNvVideoCodec at 2.2.3 in the runtime images and moved the in-tree FFmpeg build to 9.0.1 (#14925).
- Frontend aiohttp: Raised aiohttp in the Frontend image to 3.14.4, matching the other images (#15676).
Documentation
- Migration After Shutdown Grace: Clarified across the fault-tolerance guides that requests can finish during the shutdown grace period, that unfinished requests are cancelled when it expires, and that they can migrate to another worker when migration is enabled and the request allows it, with configuration guidance for rolling upgrades and scale-downs (#14872).
- Stale Documentation Links: Repaired documentation links that pointed at moved pages on
mainand returned 404 (#15308). - Agent Skills Tab: Moved Agent Skills from a single Reference page into its own documentation tab, with an overview and one page per skill category, and listed all 27 skills (#14146). The old Reference URL redirects to the new overview.
Key Dependencies
Backend runtime versions are unchanged from v1.5.0:
| Dynamo | SGLang | TensorRT-LLM | vLLM | NIXL | UCX |
|---|---|---|---|---|---|
| v1.5.1 | v0.5.18 |
v1.3.0rc25 |
v0.28.0 |
v1.4.0 (SGLang) / v1.3.1 (TensorRT-LLM) / v1.3.2 (vLLM) |
v1.21.0 |
CUDA Variants
| SGLang | TensorRT-LLM | vLLM |
|---|---|---|
| 13.0 | 13.1 | 13.0 |
Dynamo Ecosystem
| AIPerf | AISimulate | ModelExpress | vLLM-Omni |
|---|---|---|---|
v0.12.0 |
v0.12.0 |
v0.6.0 (from v0.5.0) |
v0.28.0rc1 |
Known Issues
- SGLang Sidecar Inference 500 Errors: Every inference request sent to the SGLang sidecar still fails with HTTP 500 and the error
'GenerateReqInput' object has no attribute 'batch_size'. The cause is a protocol mismatch with the bundled SGLang version that is corrected in upstream SGLang v0.5.19. No workaround is available yet; the fix is targeted for v1.6.0. - Video Decode Worker Crash: Some H.264 videos sent as multimodal input crash the worker serving the request when PyNvVideoCodec 2.2.3 decodes the final frame, and later requests to that worker fail until it restarts. The issue also affects v1.5.0 on SGLang, TensorRT-LLM, and vLLM. Image and audio inputs are not affected. No workaround is available yet; the fix is targeted for v1.6.0.
- Carried forward from v1.5.0: The other known issues documented in the v1.5.0 release notes still apply.