Skip to content

ramjet v0.6.0

Choose a tag to compare

@nessie993 nessie993 released this 13 Sep 14:12
· 98 commits to main since this release
8206b29

Ramjet v0.6.0 adds heterogeneous multi-model serving, GLM-5.3-Flash qualification, and model-aware machine-view telemetry.

Highlights:

  • Explicit model ownership per upstream, model-scoped routing, fail-closed unknown-model handling, and a combined deduplicated /v1/models response.
  • Per-model prompt, cached-prompt, and completion-token usage plus an authoritative topology view showing model, TP size, and GPU group for every engine.
  • Qualified node06 topology: Qwen3.8-Flash-Next TP4 on GPUs 0-3 and two GLM-5.3-Flash TP2 replicas on GPUs 4-5 and 6-7.
  • Immutable NVIDIA vLLM and SM120 SGLang mixed-quant GLM recipes, parser/argument/model gates, guarded canaries, backend-neutral prefill/decode metrics, and explicit DFlash2 rejection evidence.
  • Default-off time-decayed prefix affinity, hardened route-journal archival, exact-route qualification fixes, and Compose v5 rollback preservation.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.6.0@sha256:25949fb1bb97bc5c066e4a53a703d5453a49fc1aaa773f1c39e2c465c85ba976
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.6.0@sha256:7b452761667eec6f0f59f60394cda8962b43d0c70798370f2ee849819821e55a

The exact release load-balancer bytes were qualified on node06 before tagging: 3/3 upstreams healthy, combined Qwen/GLM model discovery, Qwen owner isolation, traffic observed on both GLM TP2 owners, per-upstream readiness metrics at 1, no engine restart/OOM, no recent fatal LB logs, and guarded intake at or below 43C.

Immediate rollback target retained on node06: ghcr.io/helixml/ramjet:rust-9b4372d@sha256:28057c7269b61e483e57bc4333e7de80c5eb84072eb80b11879971e2b8b7e155 (ds4-loadbalancer-rollback-20260913T140903Z).