Skip to content

Releases: helixml/ramjet

ramjet v0.7.0

Choose a tag to compare

@nessie993 nessie993 released this 29 Sep 09:08
2e124d8

Ramjet v0.7.0 routes one fleet across many nodes, and across the data-parallel attention ranks inside each node.

Highlights:

  • RJ_ROUTE_AFFINITY_BASIS=relative. Each replica is scored against the warmest serving peer of its model.
    • marginal (#294) used the least-warm peer as its floor. Beyond two replicas that peer is usually cold, so marginal fell back to absolute.
    • In the fleet simulation at 40 replicas, relative kept 90% of agent turns on the replica holding their session, against 51% for absolute and marginal.
    • At two replicas it makes the same decisions as marginal.
  • RJ_TOPOLOGY_FILE. The fleet is described as named nodes and their replicas instead of index-aligned comma lists. Node names appear on ramjet_upstream_info{upstream,node} and in /health.
  • SGLang data-parallel attention ranks as upstreams (RJ_UPSTREAM_DP_RANKS, or "dp_ranks": N on a topology replica). Requests are pinned with routed_dp_rank, so prefix affinity chooses the rank. On GLM-5.3 DP8 (FP8 KV, 8×H200) with 16 agent developers, this gave 71.9 turns/min at 92.9% cached prompt, against 43.0 turns/min at 65.6% for SGLang's own round robin.
  • Fleet resilience:
    • Failover tries other nodes before a failed replica's siblings, within RJ_ROUTE_MAX_ATTEMPTS.
    • A refused connection fences all of that engine's ranks.
    • RJ_UPSTREAM_RANK_PROBE=on detects a wedged DP rank that /health still reports as up.
    • RJ_UPSTREAM_CONNECT_TIMEOUT_MS sets the connect budget, and healthy upstreams are probed concurrently.
  • Pooled upstream connections now expire after 4s idle, below the engines' 5s keep-alive (#296). This is the one changed default.
  • bench/agent_swarm_bench.py, an open-loop coding-agent fleet generator for routing A/B tests (#295).

Every other new behaviour is opt-in; existing deployments route as before. docs/multi-node.md describes the multi-node configuration.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.7.0@sha256:562b63d16253cb1cc9f823fe41bf7146c448d3de806f49db50f5421cc4342880

Qualification. This release was tagged after the full CI pipeline passed. The node06 acceptance step in RELEASE.md was not performed. The routing features were exercised on 8×H200 GLM-5.3 test nodes, not on node06.

Rollback: the last node06-qualified load balancer is ghcr.io/helixml/ramjet:v0.6.2@sha256:53047a816c8ae1dbade7e27b86d2e3eda40cfeaa48826ae06bf2ac1d35a27cc6.

🤖 Generated with Claude Code

v0.6.2

Choose a tag to compare

@nessie993 nessie993 released this 25 Sep 15:23
962b7b2

Ramjet v0.6.2 adds an opt-in long-prompt lane and ships the GLM-5.3 prefix-cache fixes for node06.

Highlights:

  • Long-prompt lane. RJ_ROUTE_LONG_PROMPT_BYTES and RJ_ROUTE_LONG_PROMPT_UPSTREAMS (lane/- per upstream) confine requests at or above a body-size threshold to designated replicas, so one very long prefill can't evict a peer's prefix cache or stall its requests. If no lane member is serving, the request routes normally. Adds ramjet_route_long_prompt_total{upstream,outcome} and route-journal v12. Unset or 0 turns it off.
  • GLM-5.3 prefix cache (node06 deployment). The hybrid model's 28-slot linear-attention state pool, not KV space, was evicting every other session's prefix whenever a ~310k-token prompt arrived. The replicas now cap cached states per radix path and add a HiCache host tier:
    • six 20k-token sessions stay 99.6% cached across a 308k-token prompt (previously 0/6);
    • twelve sessions fit on the device (previously none);
    • host restores take 0.43 s instead of 3.4 s.
  • GLM-5.3 SwiGLU clamp. The derived SGLang image routes swiglu_limit = 10.0 into the SM120 W4A16 routed-expert kernel; the pinned SGLang and FlashInfer both dropped it. Full GSM8K went from 96.29% to 96.44%, and the agent corpus passes 5/5.
  • System One API routing (RJ_UPSTREAM_APIS, /v1/systemone) and the Kev runtime recipe.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.6.2@sha256:53047a816c8ae1dbade7e27b86d2e3eda40cfeaa48826ae06bf2ac1d35a27cc6
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.6.2@sha256:5c7ee804d5d8f5032335cc0184b9e06212d7eb1779b508b0d29f8f0124c91864

These exact load-balancer bytes were qualified on node06 before tagging:

  • the guarded canary rollout passed, with 4/4 upstreams healthy and both API families exercised;
  • a 654 KB request was routed through the lane to glm53sm120-c and a short request went to glm53sm120-b;
  • no LB errors or 5xx;
  • a Helix end-to-end run of the Dubai broker intake suite (11 cases) on warm caches reached a 93.1% prompt-cache hit ratio, a 98.9% hit on each session's first large call, and a p50 TTFT of 1.10 s.

Immediate rollback retained on node06: ghcr.io/helixml/ramjet:systemone-6b025f7-8fa31d5cec1d@sha256:bcd132f97171f6203715e04f0fc5ab3784a8b8e8cb656e7bf74c1c39b9884f93 (ds4-loadbalancer-rollback-20260925T151249Z). The lane alone can be turned off with RJ_ROUTE_LONG_PROMPT_BYTES=0.

ramjet v0.6.1

Choose a tag to compare

@nessie993 nessie993 released this 15 Sep 18:58
e90ccc3

Ramjet v0.6.1 fixes impossible negative cache-hit percentages in machine view after a busy interval ages into an idle window.

Highlights:

  • Empty rolling windows now clear floating-point accumulator residue and render as absent instead of publishing a negative percentage.
  • Added a focused frontend regression that reproduces the cancellation failure.
  • Pull-request and release pipelines now test and build the machine-view UI.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.6.1@sha256:bdd58ba1d2d98240df76f85b595a4d52420feedf809dda6462b36f97ce14f0a1
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.6.1@sha256:250c9dd79331fa357e5d85f5a93d886b8efd6b93e45ed9fc1bc416633f5acc7d

The exact release load-balancer bytes were qualified on node06 before tagging: 3/3 upstreams healthy, both model IDs present, Qwen owner isolation and both GLM owners exercised by the guarded rollout, no LB warnings/errors, and maximum observed chassis intake of 45C. The production 1-hour/400-point cache series rendered within 0–99.94% with zero out-of-range points.

Immediate rollback retained on node06: ghcr.io/helixml/ramjet:rust-8206b29@sha256:25949fb1bb97bc5c066e4a53a703d5453a49fc1aaa773f1c39e2c465c85ba976 (ds4-loadbalancer-rollback-20260915T185341Z).

ramjet v0.6.0

Choose a tag to compare

@nessie993 nessie993 released this 13 Sep 14:12
8206b29

Ramjet v0.6.0 adds heterogeneous multi-model serving, GLM-5.3-Flash qualification, and model-aware machine-view telemetry.

Highlights:

  • Explicit model ownership per upstream, model-scoped routing, fail-closed unknown-model handling, and a combined deduplicated /v1/models response.
  • Per-model prompt, cached-prompt, and completion-token usage plus an authoritative topology view showing model, TP size, and GPU group for every engine.
  • Qualified node06 topology: Qwen3.8-Flash-Next TP4 on GPUs 0-3 and two GLM-5.3-Flash TP2 replicas on GPUs 4-5 and 6-7.
  • Immutable NVIDIA vLLM and SM120 SGLang mixed-quant GLM recipes, parser/argument/model gates, guarded canaries, backend-neutral prefill/decode metrics, and explicit DFlash2 rejection evidence.
  • Default-off time-decayed prefix affinity, hardened route-journal archival, exact-route qualification fixes, and Compose v5 rollback preservation.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.6.0@sha256:25949fb1bb97bc5c066e4a53a703d5453a49fc1aaa773f1c39e2c465c85ba976
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.6.0@sha256:7b452761667eec6f0f59f60394cda8962b43d0c70798370f2ee849819821e55a

The exact release load-balancer bytes were qualified on node06 before tagging: 3/3 upstreams healthy, combined Qwen/GLM model discovery, Qwen owner isolation, traffic observed on both GLM TP2 owners, per-upstream readiness metrics at 1, no engine restart/OOM, no recent fatal LB logs, and guarded intake at or below 43C.

Immediate rollback target retained on node06: ghcr.io/helixml/ramjet:rust-9b4372d@sha256:28057c7269b61e483e57bc4333e7de80c5eb84072eb80b11879971e2b8b7e155 (ds4-loadbalancer-rollback-20260913T140903Z).

ramjet v0.5.0

Choose a tag to compare

@nessie993 nessie993 released this 01 Sep 20:53
cd85aa3

Ramjet v0.5.0 adds adaptive engine topology, authenticated topology control and audit/recovery surfaces, and a 15-second trailing average for GPU utilization in machine view. Missing host-agent samples remain unavailable instead of appearing as zero.

Immutable images:

  • Load balancer: ghcr.io/helixml/ramjet:v0.5.0@sha256:c3fc5723a0dba51f9bb8eced77648cf0b05788039e90fc638fbd8c19adec70d8
  • Snapshot companion: ghcr.io/helixml/ramjet:companion-v0.5.0@sha256:d6504e2a1cad6dc3a611084c0c2143ccb90d8025c372b59d3d91aac036e31d13

The exact load-balancer bytes were qualified on node06 before tagging: 2/2 TP4 admission, authenticated guarded inference smoke, dashboard/UI healthy, no recent fatal logs, and no engine restart. The measured load-balancer interruption upper bound was 4 seconds.

Rollback on node06 retains the prior local image ID sha256:d4d88dc98442d5da5150e85e29b2b7a90527faa8bc1424409e3216729a895768.

ramjet 0.4.0

Choose a tag to compare

@nessie993 nessie993 released this 20 Aug 13:17
1d0223e

Idle drain grows an actuator

  • RJ_IDLE_DRAIN_ACTUATOR=sleep lets the LB carry out its own park decision through vLLM sleep mode (POST /sleep / POST /wake_up). Actuation is gated on drain mode; observe remains consequence-free. A parked or waking replica stays fenced from routing, because a sleeping vLLM engine hangs rather than refuses.
  • RJ_IDLE_DRAIN_RELEASE=utilization releases an individually quiet replica while its peers serve. RJ_IDLE_DRAIN_MAX_PARKED bounds host memory: level-1 sleep does not return offloaded weights on wake.

Serving recipes

  • deploy/qwen38_27b/ documents two qualified stacks: the vLLM FP8+MTP topology family (full feature surface) and the SGLang NVFP4+DFlash2 overlay (eight single-GPU engines, fastest single-stream decode). Both serve the same model name. The SGLang tool-call parser must be qwen3_coder — the qwen name is the Qwen2.5 JSON detector and silently swallows Qwen3.8's XML tool calls.

Machine view

  • Gen tok/s tile shows a 30s mean instead of a 30s max (completion-tick spikes read as rates the fleet never sustained).
  • New per-stream decode quantiles (stream_tps_p50/stream_tps_p05) from the per-request ramjet_decode_tokens_per_second histogram, with a Stream tok/s tile.
  • Tile sparkline hover is confined to the chart's own bounds.

Images

  • ghcr.io/helixml/ramjet:v0.4.0@sha256:467e7edf40c8fcad29e741cbba52ca571cbae0261d94cff008aa6bcdb737ea1b
  • ghcr.io/helixml/ramjet:companion-v0.4.0@sha256:6d00646e40c0a3fed78b8a33d8136e52a0c46f0d5287c84bca00e61f22474d34

Qualified on node06 (build 513 → deployed by digest → tag build 514 promoted with digest equality). Rollback target: ghcr.io/helixml/ramjet:v0.3.0@sha256:2489110adbed61b7a2a415784d817d48803c234c97509fd2e0eb180c5e7453b3.

🤖 Generated with Claude Code

ramjet 0.3.0

Choose a tag to compare

@nessie993 nessie993 released this 18 Aug 13:50
37f5286

Exports every metric under the ramjet_ prefix. The ds4proxy_ prefix had outlived two project renames.

See CHANGELOG.md.

Breaking

Every metric that began ds4proxy_ now begins ramjet_. Names, labels, and types are otherwise unchanged.

Prometheus has no history under the new names. A panel or alert whose window spans the upgrade shows a gap rather than a join, and anything still querying ds4proxy_* stops returning data the moment the new binary starts. There is no compatibility shim: dual-emitting both prefixes would have doubled a 258KB scrape payload polled at 1s, and the gap was judged the cheaper cost.

Update any external dashboard, alert rule, recording rule, or script that names a metric. The canonical Grafana dashboard ships updated in this release.

Images

Load balancer ghcr.io/helixml/ramjet:v0.3.0@sha256:2489110adbed61b7a2a415784d817d48803c234c97509fd2e0eb180c5e7453b3
Companion ghcr.io/helixml/ramjet:companion-v0.3.0@sha256:776a988016d970842e1c40d211da0918237c35a96da62c760dbddae35e8bd413
Rollback target ghcr.io/helixml/ramjet:v0.2.0@sha256:9a7e2a4bc131911beec428393f0a6ccab8fb7f634217e6baffa574483c762dd8

Both promoted tags resolve to the same manifests as the qualified rust-37f5286 builds, labelled org.opencontainers.image.version=0.3.0, revision 37f52861fb4145dea2cde6243b8c03e184b4f726.

Qualification

572 Rust tests, 477 Python tests, Drone quality plus release pipeline green.

Deployed to node06 (qwen38_27b, 4× vLLM TP2 on 8× RTX PRO 6000) as an LB-only recreate: 5s, engines untouched (image, start time and restart count identical before and after), 4/4 upstreams healthy, /health ok.

The rename verified on the box after the roll: 1096 ramjet_ series and zero ds4proxy_. The compose render was diffed baseline-against-candidate first, with the file list derived from the container's com.docker.compose.project.config_files label; the only difference was the image line.

Request-generating benchmark gates were not run: node06 carries live production traffic under the cooling moratorium, so acceptance rests on observed production traffic rather than a synthetic matrix.

ramjet 0.2.0

Choose a tag to compare

@nessie993 nessie993 released this 18 Aug 11:31
68fb5dc

Renames the project to ramjet, adds the machine-view dashboard and multi-model serving, and makes the cache-hit ratio reportable against engines that never return prompt_tokens_details.cached_tokens.

See CHANGELOG.md for the full entry.

Images

Load balancer ghcr.io/helixml/ramjet:v0.2.0@sha256:9a7e2a4bc131911beec428393f0a6ccab8fb7f634217e6baffa574483c762dd8
Companion ghcr.io/helixml/ramjet:companion-v0.2.0@sha256:d9411d99be1bc5b78a55a2cdbd4d67490e7822a3ce10e99e5f49555c7f49b357
Rollback target ghcr.io/helixml/ramjet:rust-29d9e92@sha256:e303ed167f2723cec0e2d6f57b027cebfc8efd5f78ecc4b6831d9c31994978a8

Both promoted tags resolve to the same manifests as the qualified rust-68fb5dc / companion-rust-68fb5dc builds, and carry org.opencontainers.image.version=0.2.0 with revision 68fb5dc0111c65c761c040fa0c4ec7beb4c431cc.

Breaking

  • Environment prefix MD_ → RJ_; headers are now X-Ramjet-*; benchmark prefix MINI_DYNAMO_ → RAMJET_.
  • Images publish as ghcr.io/helixml/ramjet.
  • build_engine_sample returns EngineScrape; take .sample for the published shape.

The ds4proxy_ metric prefix is deliberately unchanged for dashboard continuity.

Qualification

572 Rust tests, 476 Python tests, Drone quality plus release pipeline green.

Deployed to node06 (qwen38_27b, 4× vLLM TP2 on 8× RTX PRO 6000) as an LB-only recreate with engines untouched: 4/4 upstreams healthy, /health ok, 2.9s recreate.

The cache-hit fix verified live — the ratio went from permanently absent to 98.46% sourced from engine_prefix_cache, cross-checked against the engines' own counters at 94.4–96.0% per active interval.

Request-generating benchmark gates were not run: node06 carries live production traffic under the cooling moratorium, so acceptance rests on the existing production traffic rather than a synthetic matrix.

mini-dynamo v0.1.0

Choose a tag to compare

@nessie993 nessie993 released this 13 Aug 13:34
b0e0700

mini-dynamo's first public Rust release: an OpenAI-compatible, KV-cache-locality-aware inference router for multi-replica serving.

Stable serving surface:

  • approximate prefix-locality plus weighted-load routing;
  • health-gated failover and replica-aware /health readiness;
  • immediate upstream cancellation when the client disconnects;
  • streaming compatibility shims, usage/TTFT/cache metrics, and privacy-bounded route journals;
  • bounded local/remote tokenizer observation with unconditional approximate fallback.

Exact KV events, compact snapshot companions, and exact placement remain experimental, non-mutating shadow or disabled by default.

Qualified on node06 (8× RTX PRO 6000, two DSpark/vLLM TP4 replicas):

  • 2/2 healthy; zero LB or engine restarts;
  • 206/206 successful real-workflow chat requests, no upstream error series or warning/error logs;
  • cancellation released the LB reservation 46ms after client close and vLLM work 269ms after close;
  • 82.5% cached prompt tokens in the fresh 3-app × 4-session × 2-turn locality gate;
  • c12 shared-app: exact 6/6 split, 12/12 success, 452 tok/s;
  • isolated TP4 c12/max256: 794.5 tok/s while production remained active on the peer.

Immutable images:

  • LB: ghcr.io/helixml/mini-dynamo:v0.1.0@sha256:62d949e0e6b3880796fab6c12f148f24d3f76449cb8397da6e81fe6e57dd70a1
  • Companion: ghcr.io/helixml/mini-dynamo:companion-v0.1.0@sha256:4af08be5c011ac56d1bde2463e525c1d57d9ddd21391a3565ec55183566d9f95
  • node06 rollback LB: ghcr.io/helixml/mini-dynamo:rust-r33-replay-profile-afdd3ed@sha256:26f7a30fb5523be5b8fdecc251545a33580eb9b4fb8c66eba4b512de7a32052f

Both release images preserve OCI version 0.1.0 and revision b0e070073d4266018d2f907ff35a7ee88adfdcd4. The semver manifests were copied from the exact qualified SHA-tagged images without rebuilding.