Skip to content

v0.9.0

Latest

Choose a tag to compare

@sie-release-automation sie-release-automation released this 30 Sep 13:19
2337fad

0.9.0 (2026-09-30)

⚠ BREAKING CHANGES

  • helm: nats.auth.enabled defaults to true. Upgrading restarts NATS and rolls sie-config, the gateway, and the workers; pods that have not rolled yet are refused until they do, and memory-backed queued work is lost as on any NATS restart. To avoid the gap, upgrade once with nats.auth.allowAnonymous=true, then again without it. helm upgrade --reuse-values now fails the render because the reused NATS values lack the server wiring; use --reset-then-reuse-values or -f. With an external NATS server (nats.install=false), set nats.auth.existingSecrets.{config,gateway,worker} and create the users, or set nats.auth.enabled=false. nats-box and the NATS helm test pod are disabled by default. Credentials embedded in SIE_NATS_URL are not used.
  • helm: values-aws.yaml, values-gke.yaml, and values-aks.yaml no longer enable the gateway Ingress, so a plain upgrade with them removes the existing host-less, TLS-less Ingress. A gateway Ingress now renders only with gateway auth (gateway.auth.mode=static with gateway.auth.tokenSecretName), the oauth2-proxy edge on ingress-nginx, or ingress.allowUnauthenticated=true, and only with TLS or ingress.allowPlaintext=true; set both opt-ins to keep the previous catch-all Ingress. A LoadBalancer or NodePort gateway Service needs gateway auth or gateway.service.allowUnauthenticated=true. POST /v1/pools now rejects a warm floor above SIE_GATEWAY_POOL_MAX_MINIMUM_WORKER_COUNT (default 4) or a TTL above SIE_GATEWAY_POOL_MAX_TTL_S (default 3600) with 400, a pool beyond SIE_GATEWAY_MAX_POOLS (default 64) or named default with 403. The Python and TypeScript SDKs now send SIE_API_KEY when no API key is passed and the base URL has the same origin as SIE_BASE_URL; pass an empty API key to opt out.
  • config: the gateway no longer receives the sie-config admin token, so with gateway auth enabled its admin routes (POST, PUT and DELETE under /v1/pools, /v1/admin and /v1/configs) answer 403 until gateway.auth.adminTokenSecretName is set. sie-config, gateway and worker sidecar images older than this chart do not work with the split tokens.

Features

  • examples: add document-to-markdown-olmocr, LightOnOCR-2-1B on all of olmOCR-Bench (#441) (934a2a6)
  • examples: add support-assistant-policy, the /chat support-rules run (#438) (533ed0f)
  • generation: report prefix-cache hits as usage.prompt_tokens_details.cached_tokens (#386) (c125504)
  • helm: authenticate every NATS connection with per-component users (#421) (60f30ab)
  • models: add tencent/Hy-MT2-1.8B (#360) (08c111f)
  • models: add thinking profiles for Qwen3.8-27B-FP8 (#383) (ad87d96)
  • score: report the caller's content tokens separately from the reranker prompt template (#442) (b17a906)
  • server: add operator-defined upstreams with credential and egress controls (#431) (fe22241)
  • server: serve GLiClass multilang-ultra and the layer-wise v1.0 checkpoints (#420) (5699f3e)
  • server: serve TopK-Embed-V1 multi-vector models (0.8B and 2B) (#417) (36c6129)

Bug Fixes

  • align deadline dashboards and workspace resolver contracts (#412) (a0805c3)
  • config: give gateways and worker sidecars a read-only sie-config token (#418) (ef3eb69)
  • config: validate model config writes against the worker schema and support chart rollback (#392) (296998f)
  • gateway: carry request deadlines to workers and bound direct generation (#401) (8ec3b7e)
  • gateway: distinguish request body read failures from size limits (#385) (d7e0ffa)
  • gateway: fence configuration exports against concurrent updates (#387) (1b83d3e)
  • gateway: keep other bundle models routable while workers predate a new adapter (#419) (a6eb6bb)
  • generation: verify strict structured output and serve grammars on grammar-safe profiles (#399) (5e2ae50)
  • helm: generate the sie-config admin token and gate gateway readiness (#396) (7060f20)
  • helm: require authenticated, TLS-protected gateway exposure and bound the pool API (#393) (bc2e66a)
  • integrations: send explicit item ids so rerankers map scores back (#410) (c07db06)
  • keep out-of-range timeouts from panicking the gateway and sidecar (#411) (658d1d4)
  • models: serve Iso-ModernColBERT with its PyLate recipe (#436) (422232f)
  • preserve trace context across local ingest handoffs (#405) (1e12c2d)
  • runtime: preserve model retries and parked shutdown settlement (#408) (7d70326)
  • runtime: recover MLX exits and isolate invalid model configs (#413) (51699dc)
  • sdk: retry only requests that never reached the server (#400) (0860d41)
  • server: fail only the over-long GLiClass item under overflow_policy error (#429) (0db8a34)
  • server: keep loaded models serving while another model loads or is evicted (#398) (88ee59a)
  • server: pause GLiClass graph recording after an out-of-memory first replay (#434) (1f5426d)
  • server: refuse GLiClass items whose labels leave no room for the document (#427) (b5778bd)
  • server: reject token ids and apply model profiles on /v1/embeddings (#395) (078c6e5)
  • server: retry transient model-load failures and reload exited engines (#394) (8767c81)
  • server: send float16 multivectors to the sidecar as bytes (#416) (7ee9f75)
  • telemetry: count sidecar barrier NAKs of unsupported models as model_unsupported (#432) (a009826)
  • telemetry: preserve long finite lifecycle durations (#407) (0edcfb3)
  • telemetry: preserve safe per-request batch timing (#402) (f0cb865)
  • telemetry: raise the sidecar NATS series budget for the deadline reason (#428) (db02e21)
  • telemetry: rebuild remote metric scalar attributes (#406) (91079ac)
  • telemetry: record failures and streamed response lifecycle (#404) (e4e7ec0)
  • telemetry: validate exported log resource and completion fields (#409) (ed0ad32)

Performance Improvements

  • models: load five more DeBERTa-v3 GLiClass models with bucketed CUDA graphs (#422) (aa9b323)
  • models: raise GLiClass windows to the reference 1,024 tokens (#425) (7c89f6b)
  • models: serve MADLAD from its bfloat16 CTranslate2 artifact (#433) (7256a68)
  • server: bound GLiClass CUDA graph keys so mixed traffic keeps replaying (#389) (5c50c33)
  • server: fuse the ModernBERT flash RoPE for models as accurate against float32 (#424) (f18d6d5)
  • server: replay ModernBERT flash forwards as CUDA graphs (#423) (4c4c3a2)
  • server: run GLiClass ModernBERT encoders through the flash-attention varlen stack (#414) (77577f9)