Skip to content

AINode 0.5.20

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 17 Sep 21:19
· 67 commits to main since this release
b0568bc

Engine activity as the liveness signal, ranks kept in step through autotune, OpenCode in the harness bench, catalog shapes in every record.

Added

  • ainode-bench harness --claude-effort <level> sets the reasoning effort Claude Code runs with (#127). Claude Code sends effort "high" by default, and a served chat template does not have to accept that value: Qwen3.8-Flash-Next takes only xhigh, medium and low, so every request came back API Error: 400 Unexpected reasoning effort high, the harness crashed in about 0.3 s per attempt and the model scored 0/10 on a suite the other three harnesses were passing. The same ten tasks then scored 8/10 and 10/10 at medium. The level is appended to the claude argv as --effort <level> and nothing else about the invocation moves; unset stays the default and sends nothing, so every number already recorded (Ornith, Qwen3.8 27B) keeps its meaning. It is written down where a reader will find it: the run's settings as claude_effort and the claude block's options as effort, both absent when the flag was not used, because a run with the agent's own default is a different statement from a null. Docs: the claude section of bench/harness/README.md.

Fixed

  • The harness bench record's placement.node now names the node that serves the model, not the master the CLI talked to. ainode-bench harness built the placement block by calling describe_via_http with the --ainode web base as both the control-plane and engine URL, so when DeepSeek V4 Flash was served on Spark-2 while the CLI pointed at Spark-1, the record's placement.node read "Spark-1". The CLI now resolves the serving node through the master's /api/server/status loaded_models (the same fleet view the speed bench and the in-product bench use), calls describe_via_http against that node's own web base, and records the engine port in placement. A model served on the master is unchanged; an unloaded model produces the same "does not report serving" warning.
  • The harness bench's OpenCode adapter now runs in the task's directory instead of the one the bench was started from, which is the whole of #118. Every --harness opencode invocation from the runner exited 1 in about 1 s with {"type":"error",...,"message":"Unexpected server error. Check server logs for details."} on stdout and an empty log, while the identical command run by hand passed, so the adapter was carried as unverified and the DeepSeek harness record's OpenCode rows are marked not measured. --print-logs --log-level DEBUG off a runner-shaped invocation named it in one run: opencode creates its instance in the working directory, reads the run's project-local opencode.json there, and then creates a second instance and builds the session against the directory named by PWD - which a subprocess started with cwd= still inherits from its parent, so it was the shell the bench was launched in. That instance has no ainode-bench provider, so the session died with ProviderModelNotFoundError: Model not found: ainode-bench/<model id>, reported on stdout as the generic server error. By hand it passed because a person cds in first and PWD is then correct; that asymmetry, not flakiness, is why it looked intermittent, and it is why the other four harnesses were never affected (none of them reads PWD to find its project). Two fixes, one cause: the argv now states --dir <workdir>, and _launch sets PWD to the launch cwd and drops a stale OLDPWD for every harness, since the wrong value also reaches whatever shell an agent spawns for its own tools. Verified by the runner itself, --only-tasks isogram against Ornith 1.5 35B A3B NVFP4 on the fleet: three consecutive runs, pass@1 1/1 at 25.5 s, 26.7 s and 26.5 s, zero crashes, where the same command on the previous commit crashed in 0.9 s twice out of two.
  • An OpenCode bench run no longer reads the operator's global config, and a failed one now says what failed. OPENCODE_CONFIG_DIR overrides the XDG config search outright, so the four isolated XDG_* variables the adapter already set were bypassed by any exported value (Orca exports one) and somebody's global config, plugins and agents were inside the measurement; the overlay now points that variable at the run's own empty config directory too, the same rule DSH_HOME and CLAUDE_CONFIG_DIR follow. The argv also carries --print-logs --log-level ERROR, which is diagnostics rather than behaviour - the logs go to stderr, stdout stays the NDJSON event stream the adapter parses - so a future failure reaches the record's stderr_tail with a named cause instead of the opaque "Unexpected server error" that #118 had to work from; verified against the old broken shape, where it prints the ProviderModelNotFoundError in 1388 bytes.
  • The bind wait watches what the engine is DOING; log silence is only the second opinion (#112). vLLM prints nothing at all while it loads weight shards, runs torch.compile and autotunes FlashInfer, so "no log line for N seconds" could never tell a wedged engine from a working one: three healthy starts in two days were declared silent (Qwen3.8 27B NVFP4 quiet for 206 s between shard updates, Nemotron 3.5 Lightning for 363 s inside autotune and graph capture, then a 48-minute autotune), and the window had to grow 120 to 360 to 900 s to keep up. Since 0.5.13 that verdict broke nothing (the relaunch is a no-op while the container is alive) but it spent the engine's single retry and wrote a wrong reason into the log. Backends now publish activity_mark(): for NvidiaBackend the last time THAT engine's container was seen burning CPU, read from the container's cgroup cpu.stat when the tree is visible and otherwise from docker stats --no-stream, cached so one bind poll costs one probe and run off the event loop thread. A log line OR an advancing activity mark resets the silence clock, and the start is called dead only when both have been quiet for the whole budget, which is why the budget comes back down to 300 s (engine_bind_log_silence_seconds): a truly wedged engine does no work at all and is caught in five minutes again instead of fifteen. The hard death signal (container exited) and the 1800 s ceiling are unchanged, and a backend that cannot see its engine's container reports None and keeps the old log-driven behaviour. GPU utilization is deliberately not part of this: the pynvml path reads device 0 as a whole, so it cannot say which stacked engine is busy, and it reads 0% during exactly the phases that need covering.
  • engine_ready in /api/nodes means the engine answers /v1/models, not that the node is reachable (the second half of #112). The field was derived from status, which is heartbeat health (online/stale/offline) and true of every node whose orchestrator is up, because ClusterNode.from_discovered replaces the announced engine state with that health: Spark-4 was published engine_ready: true while its engine was still loading and its loaded_models was empty. ClusterNode now carries the node's own word on its engine (engine_status, straight from the announcement, which that node's sync loop already sets from a live /v1/models probe), the local row answers from a fresh localhost probe rather than the ready latch, and a member node (which runs no local vLLM) still reads ready once discovered. status is untouched, so the topology's online count keeps working.
  • The two ranks of an mp launch stay in step through FlashInfer autotune, instead of gloo killing the pair 48 minutes in (#134). Kernel warmup runs per rank. On Qwen3.8-Flash-Next the head, warm from an earlier launch, finished autotune in tens of minutes while the peer tuned 856 profiles from cold, sat 1800 s on its first trtllm::fused_moe::gemm1 profile, and the barrier after autotune died with Connection closed by peer on one side and Application timeout caused pair closure on the other. 1800 s is not a kernel hang: it is PyTorch's default timeout for the CPU (gloo) group, which is the group the barrier and the autotune sync both use. Three things change. The head now ships its compiled-kernel cache to a peer before that peer's container starts (_ensure_peer_has_jit_cache): one transfer per child of the recipe's VLLM_CACHE_ROOT, skipped per child when the peer already has it, and a transfer that fails is a warning rather than a refused launch, since a missing cache costs minutes and a refused launch costs the model. The Flash-Next entry turns autotune off with --no-enable-flashinfer-autotune (the off form of KernelConfig.enable_flashinfer_autotune, which kernel_warmup.py reads to skip the whole phase) and raises both collective floors to 5400 s, above the worst warmup measured on the pair; the stated cost is that its fused MoE and fp4 GEMMs now run FlashInfer's heuristic tactic rather than a measured one, so decode is slower than the 26.6 tok/s of the launch that got through. It also states the cache roots it was missing, so its JIT output persists per node inside the mounted HF cache instead of dying with the container and re-tuning from cold every launch. The bind wait's ceiling goes from 1800 s to 3600 s, because a launch that took 36 minutes to ready was being killed by the ceiling as well; the ceiling is the last resort and the activity signal from #112 still calls a genuinely wedged engine dead inside five minutes, so nothing waits an hour on a dead start. Worth writing down for the next person: seeding a peer's autotune file cannot help on its own. Only rank 0 reads that file and broadcasts its bytes to the other ranks, and the path holds a sha256 of the whole VllmConfig hash, so a hand-copied directory is both unread and easy to orphan. DeepSeek V4 Flash keeps its own autotune and timeouts exactly as proven.
  • Catalog entries state their active parameter count and architecture, so a bench record stops reporting a MoE as dense. ModelInfo had no field for either, and the record's model block was filled by reading the A<n>B marker off the model id, which works for Ornith-1.5-35B-A3B and not at all for an id that does not carry one: DeepSeek V4 Flash (284B total, 13B active) and Qwen3.8-Flash-Next (125B, 6B) were both written down as params_b 125/284, active_b 125/284, arch dense, which on GB10 is the one number that predicts decode speed, said wrong. ModelInfo now carries active_params_b and arch ("moe" / "dense"), every curated and fallback entry states its shape, and the model block takes both from the catalog when it has them and falls back to the id only for a model the catalog does not describe. An entry that says MoE without an active count (GLM-5.2 REAP, where pruning moves it) leaves active_b out and both renderers print 504B MoE rather than claiming dense. The catalog cache round-trips the new fields and a cache file written before they existed still loads. Applies to the throughput bench, the in-product /api/bench runs and the harness bench alike: all three build the block through ainode/bench/fleet.py.

Image: ghcr.io/getainode/ainode:0.5.20