Releases
v0.8.0
Compare
Sorry, something went wrong.
No results found
0.8.0 (2026-09-23)
⚠ BREAKING CHANGES
server: extra_launch_args may not carry --nccl-port, --host or --port, in any unambiguous abbreviation or --flag=value spelling. A profile carrying one fails to load, and the config service refuses the write that would create it, so profiles written straight through the config API need the same edit as the ones shipped in a chart. Replace --nccl-port with the adapter_options.loadtime.nccl_port option, which applies above tensor_parallel_size: 1 and is reserved before the flag is passed. Remove --host and --port: the server passes both for the HTTP listener it talks to.
server: a load-time option that no adapter constructor accepts is refused when the model loads instead of being ignored. extra_launch_args may not carry placement flags (including unambiguous abbreviations) and extra_env may not set CUDA_VISIBLE_DEVICES, NVIDIA_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES. The SGLang embedding adapter no longer accepts pooling_method. SGLang launch arguments spell the width as --tensor-parallel-size for every profile.
Features
carry inline video through the queued chat completions path (#329 ) (d0633d1 )
ci: add CI and release automation (4e110b9 )
config: add GLM-5.3-Flash across eight accelerators (#324 ) (a913d0c )
examples: add detect, image-classify and speech-to-text evaluations (#313 ) (8b4f3d4 )
examples: add five recorded task evaluations, with evidence on HuggingFace (#310 ) (7d6e11d )
examples: add image-search and visual-document-search evaluations (#314 ) (c303dfc )
examples: add ocr-two-stage, a recorded two-stage OCR evaluation (#303 ) (4202f7d )
examples: add three recorded task evaluations, with evidence on HuggingFace (#311 ) (8be1e04 )
examples: measure whether a second call earns its place on doc field extraction (#336 ) (f4521a9 )
examples: re-derive the yes-or-no figure for /structured-output (#335 ) (8333afa )
examples: score four guardrail models on the same twelve inputs (#319 ) (6794e34 )
examples: score grounded answers in chat, move the incident run to its own example (#317 ) (5cc5580 )
helm: add a tested high-availability values composition (#274 ) (0560f01 )
helm: make the worker /dev/shm size configurable per pool (#302 ) (e203682 )
models: add google/translategemma-4b-it (40ae372 )
models: add google/translategemma-4b-it (979961b )
models: add Qwen3-VL-8B-Instruct with text, image, and video input (#323 ) (92bd4bb )
server: accept inline video in local chat completions (#321 ) (0fbe99b )
server: move the CUDA 13 SGLang bundle to 0.5.20 for glm5_next (#320 ) (ac2358c )
server: parse and force GLM tool calls on the queued route (#328 ) (84b1a7f )
server: serve one model across several GPUs with tensor parallelism (#282 ) (a03f5db )
Bug Fixes
api: constrain native stream execution evidence pairs (7cfa50e )
api: restrict stream evidence to successful terminal chunks (f256d4e )
bound diagnostic decoding and preserve grammar refusal metadata (b70cca0 )
ci: serialize mise bootstrap to protect Node GPG state (538daff )
ci: serialize mise bootstrap to protect Node GPG state (996b07f )
ci: serialize mise bootstrap to protect Node GPG state (#287 ) (538daff )
config: add Qwen3.8 grammar fallback (#258 ) (1c7bbe8 )
config: keep GLM-5.3-Flash reasoning private on every route (#326 ) (1ffb19c )
document observed images in server generation usage (81aecec )
examples: record Granite Guardian through chat completions on guardrails (#316 ) (86f4e1e )
gateway: keep the served surface when an authoritative export shrinks it (#270 ) (c65aee5 )
gateway: pass the invalid_guard_verdict worker code through (#297 ) (7a8b09b )
gateway: refuse requests when the auth configuration is an error (#269 ) (8c84419 )
generate: preserve image usage and complete execution evidence (d266205 )
generate: preserve image usage and complete execution evidence (c9e9b7e )
generate: preserve image usage and complete execution evidence (#286 ) (d266205 )
generation: require explicit streamed candidate indexes (c0aca44 )
helm: make KEDA hook Job resources configurable (#268 ) (5e726f6 )
identify unsupported Outlines JSON Schema type values (729cadc )
identify unsupported Outlines JSON Schema type values (0c0408b )
identify unsupported Outlines JSON Schema type values (#289 ) (729cadc )
isolate grammar followers and type malformed backend errors (4c5e05f )
mcp: parse the Host header before trusting it for the OAuth origin (27ae3a1 )
mcp: parse the Host header before trusting it for the OAuth origin (604449a )
mcp: stop deriving the OAuth origin from X-Forwarded headers (#292 ) (f5c4451 )
mcp: validate bracketed IPv6 before advertising origins (a66f6a0 )
mcp: validate Host before parsing OAuth request origins (d9507f5 )
models: bound translategemma prompts to the documented 2K input context (ac810ad )
models: pin MADLAD to greedy sampling by default (b0f0eeb )
models: pin MADLAD to greedy sampling by default (a968874 )
models: pin the triton attention backend for translategemma (2dfe7b2 )
omit logprobs for rewritten guard verdicts (ba8c3b1 )
preserve generation progress and reject invalid guard verdicts (424f7ae )
preserve generation progress and reject invalid guard verdicts (398c673 )
preserve generation progress and reject invalid guard verdicts (#285 ) (424f7ae )
preserve grammar-safe generation defaults (#261 ) (3b16ffb )
preserve TensorRT-LLM completion text and stop alignment (#267 ) (d345bc0 )
release: format generated SDK package metadata (#339 ) (36bd49e )
sdk: declare terminal streaming execution evidence (0261183 )
sdk: declare terminal streaming execution evidence (#283 ) (0261183 )
sdk: document terminal streaming execution evidence (f3a05ba )
server: bound native encode, score and extract request bodies (#271 ) (8479f1c )
server: evict the least recently used group when two blocks cost the same (#294 ) (120059f )
server: fetch model weights outside the registry load lock (#272 ) (394df63 )
server: force complete GLM argument pairs and reject oversized GLM calls (#330 ) (079647d )
server: keep inline media payloads out of SGLang child logs (#325 ) (d8573f2 )
server: map GroundingDINO detections back onto the caller's labels (#277 ) (958414c )
server: narrow validated terminal result type (56efe5f )
server: pin cuda-tile for TensorRT-LLM (#254 ) (6484fab )
server: read Qwen2.5-style tool calls as Hermes JSON on the queued route (#332 ) (d8494f3 )
server: refuse listener flags in extra_launch_args (#304 ) (7a68262 )
server: reject failed SGLang generation terminals (829f80c )
server: reject failed SGLang generation terminals (3bc140d )
server: reject failed SGLang generation terminals (#284 ) (829f80c )
server: reject malformed backend finish metadata (2be0155 )
server: report a valueless --mm-process-config instead of raising (#333 ) (2f7ad9d )
server: require a usable startup budget for SGLang profiles (#301 ) (038a8d9 )
server: require complete generation candidates (afaab57 )
server: ship FFmpeg shared libraries in the SGLang runtime image (#322 ) (9d98037 )
server: tell SGLang when a native generate request starts inside reasoning (#327 ) (26702c7 )
server: validate video pixel and frame counts before the budget arithmetic (#337 ) (703841a )
tasks: stop expanding possibly-empty arrays bare under set -u (#278 ) (e0084c7 )
toolchain: upgrade Rust to 1.98.1 and patch rustls (#276 ) (a1f6fab )
tooling: use canonical Rust LLVM component (#256 ) (2183bc1 )
ts-sdk: decode base64 data URLs with media-type parameters (#247 ) (21289a1 )
validate complete first guard verdict evidence (514b07f )
validate unsupported template placeholders (#249 ) (0508e17 )
You can’t perform that action at this time.