You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
Fixed
Static assets can no longer go stale across an update (#77) — every /static URL in the served HTML carries ?v=<version> and /static
responses are Cache-Control: no-cache. After the 0.5.7 roll a browser kept
the 0.5.6 stylesheet on heuristic freshness and rendered the new chat
unstyled.
The model card only renders on the Chat view (#77); it was showing on
Cluster, Server, Models and Bench.
Startup replay no longer kills slow-but-healthy engines. The bind wait had
a fixed 300s window. On vllm/vllm-openai:v0.27.1 a GB10 node spends minutes
in FlashInfer fp4_gemm autotune and CUDA graph capture before the server
listens, so on Spark-1 (2026-09-13) the window expired at 5 minutes on a 27B
NVFP4 primary that bound at ~12 and a 35B-A3B stacked instance that bound at
~6. Both were relaunched from scratch, turning a 14-minute boot into 28. The
wait now treats an engine as alive while its container is up and its log is
still advancing, and relaunches only on evidence: container exited, log silent
past engine_bind_log_silence_seconds (default 120), or engine_bind_ceiling_seconds reached (default 1800). An engine that dies on
the way up still gets exactly one relaunch (0.5.5), and the ainode log now
carries the reason and how long the wait lasted.
Changed
"Made in Texas" replaces "Powered by argentos.ai" in the web and
onboarding footers, the CLI banner and status output, the bench report and
the status API (powered_by: ainode.dev) (#77).