Releases: helixml/ramjet
Release list
ramjet v0.7.0
Ramjet v0.7.0 routes one fleet across many nodes, and across the data-parallel attention ranks inside each node.
Highlights:
RJ_ROUTE_AFFINITY_BASIS=relative. Each replica is scored against the warmest serving peer of its model.marginal(#294) used the least-warm peer as its floor. Beyond two replicas that peer is usually cold, somarginalfell back toabsolute.- In the fleet simulation at 40 replicas,
relativekept 90% of agent turns on the replica holding their session, against 51% forabsoluteandmarginal. - At two replicas it makes the same decisions as
marginal.
RJ_TOPOLOGY_FILE. The fleet is described as named nodes and their replicas instead of index-aligned comma lists. Node names appear onramjet_upstream_info{upstream,node}and in/health.- SGLang data-parallel attention ranks as upstreams (
RJ_UPSTREAM_DP_RANKS, or"dp_ranks": Non a topology replica). Requests are pinned withrouted_dp_rank, so prefix affinity chooses the rank. On GLM-5.3 DP8 (FP8 KV, 8×H200) with 16 agent developers, this gave 71.9 turns/min at 92.9% cached prompt, against 43.0 turns/min at 65.6% for SGLang's own round robin. - Fleet resilience:
- Failover tries other nodes before a failed replica's siblings, within
RJ_ROUTE_MAX_ATTEMPTS. - A refused connection fences all of that engine's ranks.
RJ_UPSTREAM_RANK_PROBE=ondetects a wedged DP rank that/healthstill reports as up.RJ_UPSTREAM_CONNECT_TIMEOUT_MSsets the connect budget, and healthy upstreams are probed concurrently.
- Failover tries other nodes before a failed replica's siblings, within
- Pooled upstream connections now expire after 4s idle, below the engines' 5s keep-alive (#296). This is the one changed default.
bench/agent_swarm_bench.py, an open-loop coding-agent fleet generator for routing A/B tests (#295).
Every other new behaviour is opt-in; existing deployments route as before. docs/multi-node.md describes the multi-node configuration.
Immutable images:
- Load balancer:
ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d - Snapshot companion:
ghcr.io/helixml/ramjet:companion-v0.7.0@sha256:562b63d16253cb1cc9f823fe41bf7146c448d3de806f49db50f5421cc4342880
Qualification. This release was tagged after the full CI pipeline passed. The node06 acceptance step in RELEASE.md was not performed. The routing features were exercised on 8×H200 GLM-5.3 test nodes, not on node06.
Rollback: the last node06-qualified load balancer is ghcr.io/helixml/ramjet:v0.6.2@sha256:53047a816c8ae1dbade7e27b86d2e3eda40cfeaa48826ae06bf2ac1d35a27cc6.
🤖 Generated with Claude Code
v0.6.2
Ramjet v0.6.2 adds an opt-in long-prompt lane and ships the GLM-5.3 prefix-cache fixes for node06.
Highlights:
- Long-prompt lane.
RJ_ROUTE_LONG_PROMPT_BYTESandRJ_ROUTE_LONG_PROMPT_UPSTREAMS(lane/-per upstream) confine requests at or above a body-size threshold to designated replicas, so one very long prefill can't evict a peer's prefix cache or stall its requests. If no lane member is serving, the request routes normally. Addsramjet_route_long_prompt_total{upstream,outcome}and route-journal v12. Unset or0turns it off. - GLM-5.3 prefix cache (node06 deployment). The hybrid model's 28-slot linear-attention state pool, not KV space, was evicting every other session's prefix whenever a ~310k-token prompt arrived. The replicas now cap cached states per radix path and add a HiCache host tier:
- six 20k-token sessions stay 99.6% cached across a 308k-token prompt (previously 0/6);
- twelve sessions fit on the device (previously none);
- host restores take 0.43 s instead of 3.4 s.
- GLM-5.3 SwiGLU clamp. The derived SGLang image routes
swiglu_limit = 10.0into the SM120 W4A16 routed-expert kernel; the pinned SGLang and FlashInfer both dropped it. Full GSM8K went from 96.29% to 96.44%, and the agent corpus passes 5/5. - System One API routing (
RJ_UPSTREAM_APIS,/v1/systemone) and the Kev runtime recipe.
Immutable images:
- Load balancer:
ghcr.io/helixml/ramjet:v0.6.2@sha256:53047a816c8ae1dbade7e27b86d2e3eda40cfeaa48826ae06bf2ac1d35a27cc6 - Snapshot companion:
ghcr.io/helixml/ramjet:companion-v0.6.2@sha256:5c7ee804d5d8f5032335cc0184b9e06212d7eb1779b508b0d29f8f0124c91864
These exact load-balancer bytes were qualified on node06 before tagging:
- the guarded canary rollout passed, with 4/4 upstreams healthy and both API families exercised;
- a 654 KB request was routed through the lane to
glm53sm120-cand a short request went toglm53sm120-b; - no LB errors or 5xx;
- a Helix end-to-end run of the Dubai broker intake suite (11 cases) on warm caches reached a 93.1% prompt-cache hit ratio, a 98.9% hit on each session's first large call, and a p50 TTFT of 1.10 s.
Immediate rollback retained on node06: ghcr.io/helixml/ramjet:systemone-6b025f7-8fa31d5cec1d@sha256:bcd132f97171f6203715e04f0fc5ab3784a8b8e8cb656e7bf74c1c39b9884f93 (ds4-loadbalancer-rollback-20260925T151249Z). The lane alone can be turned off with RJ_ROUTE_LONG_PROMPT_BYTES=0.
ramjet v0.6.1
Ramjet v0.6.1 fixes impossible negative cache-hit percentages in machine view after a busy interval ages into an idle window.
Highlights:
- Empty rolling windows now clear floating-point accumulator residue and render as absent instead of publishing a negative percentage.
- Added a focused frontend regression that reproduces the cancellation failure.
- Pull-request and release pipelines now test and build the machine-view UI.
Immutable images:
- Load balancer:
ghcr.io/helixml/ramjet:v0.6.1@sha256:bdd58ba1d2d98240df76f85b595a4d52420feedf809dda6462b36f97ce14f0a1 - Snapshot companion:
ghcr.io/helixml/ramjet:companion-v0.6.1@sha256:250c9dd79331fa357e5d85f5a93d886b8efd6b93e45ed9fc1bc416633f5acc7d
The exact release load-balancer bytes were qualified on node06 before tagging: 3/3 upstreams healthy, both model IDs present, Qwen owner isolation and both GLM owners exercised by the guarded rollout, no LB warnings/errors, and maximum observed chassis intake of 45C. The production 1-hour/400-point cache series rendered within 0–99.94% with zero out-of-range points.
Immediate rollback retained on node06: ghcr.io/helixml/ramjet:rust-8206b29@sha256:25949fb1bb97bc5c066e4a53a703d5453a49fc1aaa773f1c39e2c465c85ba976 (ds4-loadbalancer-rollback-20260915T185341Z).
ramjet v0.6.0
Ramjet v0.6.0 adds heterogeneous multi-model serving, GLM-5.3-Flash qualification, and model-aware machine-view telemetry.
Highlights:
- Explicit model ownership per upstream, model-scoped routing, fail-closed unknown-model handling, and a combined deduplicated
/v1/modelsresponse. - Per-model prompt, cached-prompt, and completion-token usage plus an authoritative topology view showing model, TP size, and GPU group for every engine.
- Qualified node06 topology: Qwen3.8-Flash-Next TP4 on GPUs 0-3 and two GLM-5.3-Flash TP2 replicas on GPUs 4-5 and 6-7.
- Immutable NVIDIA vLLM and SM120 SGLang mixed-quant GLM recipes, parser/argument/model gates, guarded canaries, backend-neutral prefill/decode metrics, and explicit DFlash2 rejection evidence.
- Default-off time-decayed prefix affinity, hardened route-journal archival, exact-route qualification fixes, and Compose v5 rollback preservation.
Immutable images:
- Load balancer:
ghcr.io/helixml/ramjet:v0.6.0@sha256:25949fb1bb97bc5c066e4a53a703d5453a49fc1aaa773f1c39e2c465c85ba976 - Snapshot companion:
ghcr.io/helixml/ramjet:companion-v0.6.0@sha256:7b452761667eec6f0f59f60394cda8962b43d0c70798370f2ee849819821e55a
The exact release load-balancer bytes were qualified on node06 before tagging: 3/3 upstreams healthy, combined Qwen/GLM model discovery, Qwen owner isolation, traffic observed on both GLM TP2 owners, per-upstream readiness metrics at 1, no engine restart/OOM, no recent fatal LB logs, and guarded intake at or below 43C.
Immediate rollback target retained on node06: ghcr.io/helixml/ramjet:rust-9b4372d@sha256:28057c7269b61e483e57bc4333e7de80c5eb84072eb80b11879971e2b8b7e155 (ds4-loadbalancer-rollback-20260913T140903Z).
ramjet v0.5.0
Ramjet v0.5.0 adds adaptive engine topology, authenticated topology control and audit/recovery surfaces, and a 15-second trailing average for GPU utilization in machine view. Missing host-agent samples remain unavailable instead of appearing as zero.
Immutable images:
- Load balancer:
ghcr.io/helixml/ramjet:v0.5.0@sha256:c3fc5723a0dba51f9bb8eced77648cf0b05788039e90fc638fbd8c19adec70d8 - Snapshot companion:
ghcr.io/helixml/ramjet:companion-v0.5.0@sha256:d6504e2a1cad6dc3a611084c0c2143ccb90d8025c372b59d3d91aac036e31d13
The exact load-balancer bytes were qualified on node06 before tagging: 2/2 TP4 admission, authenticated guarded inference smoke, dashboard/UI healthy, no recent fatal logs, and no engine restart. The measured load-balancer interruption upper bound was 4 seconds.
Rollback on node06 retains the prior local image ID sha256:d4d88dc98442d5da5150e85e29b2b7a90527faa8bc1424409e3216729a895768.
ramjet 0.4.0
Idle drain grows an actuator
RJ_IDLE_DRAIN_ACTUATOR=sleeplets the LB carry out its own park decision through vLLM sleep mode (POST /sleep/POST /wake_up). Actuation is gated ondrainmode;observeremains consequence-free. A parked or waking replica stays fenced from routing, because a sleeping vLLM engine hangs rather than refuses.RJ_IDLE_DRAIN_RELEASE=utilizationreleases an individually quiet replica while its peers serve.RJ_IDLE_DRAIN_MAX_PARKEDbounds host memory: level-1 sleep does not return offloaded weights on wake.
Serving recipes
deploy/qwen38_27b/documents two qualified stacks: the vLLM FP8+MTP topology family (full feature surface) and the SGLang NVFP4+DFlash2 overlay (eight single-GPU engines, fastest single-stream decode). Both serve the same model name. The SGLang tool-call parser must beqwen3_coder— theqwenname is the Qwen2.5 JSON detector and silently swallows Qwen3.8's XML tool calls.
Machine view
- Gen tok/s tile shows a 30s mean instead of a 30s max (completion-tick spikes read as rates the fleet never sustained).
- New per-stream decode quantiles (
stream_tps_p50/stream_tps_p05) from the per-requestramjet_decode_tokens_per_secondhistogram, with a Stream tok/s tile. - Tile sparkline hover is confined to the chart's own bounds.
Images
ghcr.io/helixml/ramjet:v0.4.0@sha256:467e7edf40c8fcad29e741cbba52ca571cbae0261d94cff008aa6bcdb737ea1bghcr.io/helixml/ramjet:companion-v0.4.0@sha256:6d00646e40c0a3fed78b8a33d8136e52a0c46f0d5287c84bca00e61f22474d34
Qualified on node06 (build 513 → deployed by digest → tag build 514 promoted with digest equality). Rollback target: ghcr.io/helixml/ramjet:v0.3.0@sha256:2489110adbed61b7a2a415784d817d48803c234c97509fd2e0eb180c5e7453b3.
🤖 Generated with Claude Code
ramjet 0.3.0
Exports every metric under the ramjet_ prefix. The ds4proxy_ prefix had outlived two project renames.
See CHANGELOG.md.
Breaking
Every metric that began ds4proxy_ now begins ramjet_. Names, labels, and types are otherwise unchanged.
Prometheus has no history under the new names. A panel or alert whose window spans the upgrade shows a gap rather than a join, and anything still querying ds4proxy_* stops returning data the moment the new binary starts. There is no compatibility shim: dual-emitting both prefixes would have doubled a 258KB scrape payload polled at 1s, and the gap was judged the cheaper cost.
Update any external dashboard, alert rule, recording rule, or script that names a metric. The canonical Grafana dashboard ships updated in this release.
Images
| Load balancer | ghcr.io/helixml/ramjet:v0.3.0@sha256:2489110adbed61b7a2a415784d817d48803c234c97509fd2e0eb180c5e7453b3 |
| Companion | ghcr.io/helixml/ramjet:companion-v0.3.0@sha256:776a988016d970842e1c40d211da0918237c35a96da62c760dbddae35e8bd413 |
| Rollback target | ghcr.io/helixml/ramjet:v0.2.0@sha256:9a7e2a4bc131911beec428393f0a6ccab8fb7f634217e6baffa574483c762dd8 |
Both promoted tags resolve to the same manifests as the qualified rust-37f5286 builds, labelled org.opencontainers.image.version=0.3.0, revision 37f52861fb4145dea2cde6243b8c03e184b4f726.
Qualification
572 Rust tests, 477 Python tests, Drone quality plus release pipeline green.
Deployed to node06 (qwen38_27b, 4× vLLM TP2 on 8× RTX PRO 6000) as an LB-only recreate: 5s, engines untouched (image, start time and restart count identical before and after), 4/4 upstreams healthy, /health ok.
The rename verified on the box after the roll: 1096 ramjet_ series and zero ds4proxy_. The compose render was diffed baseline-against-candidate first, with the file list derived from the container's com.docker.compose.project.config_files label; the only difference was the image line.
Request-generating benchmark gates were not run: node06 carries live production traffic under the cooling moratorium, so acceptance rests on observed production traffic rather than a synthetic matrix.
ramjet 0.2.0
Renames the project to ramjet, adds the machine-view dashboard and multi-model serving, and makes the cache-hit ratio reportable against engines that never return prompt_tokens_details.cached_tokens.
See CHANGELOG.md for the full entry.
Images
| Load balancer | ghcr.io/helixml/ramjet:v0.2.0@sha256:9a7e2a4bc131911beec428393f0a6ccab8fb7f634217e6baffa574483c762dd8 |
| Companion | ghcr.io/helixml/ramjet:companion-v0.2.0@sha256:d9411d99be1bc5b78a55a2cdbd4d67490e7822a3ce10e99e5f49555c7f49b357 |
| Rollback target | ghcr.io/helixml/ramjet:rust-29d9e92@sha256:e303ed167f2723cec0e2d6f57b027cebfc8efd5f78ecc4b6831d9c31994978a8 |
Both promoted tags resolve to the same manifests as the qualified rust-68fb5dc / companion-rust-68fb5dc builds, and carry org.opencontainers.image.version=0.2.0 with revision 68fb5dc0111c65c761c040fa0c4ec7beb4c431cc.
Breaking
- Environment prefix
MD_→RJ_; headers are nowX-Ramjet-*; benchmark prefixMINI_DYNAMO_→RAMJET_. - Images publish as
ghcr.io/helixml/ramjet. build_engine_samplereturnsEngineScrape; take.samplefor the published shape.
The ds4proxy_ metric prefix is deliberately unchanged for dashboard continuity.
Qualification
572 Rust tests, 476 Python tests, Drone quality plus release pipeline green.
Deployed to node06 (qwen38_27b, 4× vLLM TP2 on 8× RTX PRO 6000) as an LB-only recreate with engines untouched: 4/4 upstreams healthy, /health ok, 2.9s recreate.
The cache-hit fix verified live — the ratio went from permanently absent to 98.46% sourced from engine_prefix_cache, cross-checked against the engines' own counters at 94.4–96.0% per active interval.
Request-generating benchmark gates were not run: node06 carries live production traffic under the cooling moratorium, so acceptance rests on the existing production traffic rather than a synthetic matrix.
mini-dynamo v0.1.0
mini-dynamo's first public Rust release: an OpenAI-compatible, KV-cache-locality-aware inference router for multi-replica serving.
Stable serving surface:
- approximate prefix-locality plus weighted-load routing;
- health-gated failover and replica-aware
/healthreadiness; - immediate upstream cancellation when the client disconnects;
- streaming compatibility shims, usage/TTFT/cache metrics, and privacy-bounded route journals;
- bounded local/remote tokenizer observation with unconditional approximate fallback.
Exact KV events, compact snapshot companions, and exact placement remain experimental, non-mutating shadow or disabled by default.
Qualified on node06 (8× RTX PRO 6000, two DSpark/vLLM TP4 replicas):
- 2/2 healthy; zero LB or engine restarts;
- 206/206 successful real-workflow chat requests, no upstream error series or warning/error logs;
- cancellation released the LB reservation 46ms after client close and vLLM work 269ms after close;
- 82.5% cached prompt tokens in the fresh 3-app × 4-session × 2-turn locality gate;
- c12 shared-app: exact 6/6 split, 12/12 success, 452 tok/s;
- isolated TP4 c12/max256: 794.5 tok/s while production remained active on the peer.
Immutable images:
- LB:
ghcr.io/helixml/mini-dynamo:v0.1.0@sha256:62d949e0e6b3880796fab6c12f148f24d3f76449cb8397da6e81fe6e57dd70a1 - Companion:
ghcr.io/helixml/mini-dynamo:companion-v0.1.0@sha256:4af08be5c011ac56d1bde2463e525c1d57d9ddd21391a3565ec55183566d9f95 - node06 rollback LB:
ghcr.io/helixml/mini-dynamo:rust-r33-replay-profile-afdd3ed@sha256:26f7a30fb5523be5b8fdecc251545a33580eb9b4fb8c66eba4b512de7a32052f
Both release images preserve OCI version 0.1.0 and revision b0e070073d4266018d2f907ff35a7ee88adfdcd4. The semver manifests were copied from the exact qualified SHA-tagged images without rebuilding.