Repository navigation
AINode 0.5.20
Engine activity as the liveness signal, ranks kept in step through autotune, OpenCode in the harness bench, catalog shapes in every record.
Added
ainode-bench harness --claude-effort <level>sets the reasoning effort Claude Code runs with (#127). Claude Code sends effort "high" by default, and a served chat template does not have to accept that value: Qwen3.8-Flash-Next takes only xhigh, medium and low, so every request came backAPI Error: 400 Unexpected reasoning effort high, the harness crashed in about 0.3 s per attempt and the model scored 0/10 on a suite the other three harnesses were passing. The same ten tasks then scored 8/10 and 10/10 at medium. The level is appended to the claude argv as--effort <level>and nothing else about the invocation moves; unset stays the default and sends nothing, so every number already recorded (Ornith, Qwen3.8 27B) keeps its meaning. It is written down where a reader will find it: the run'ssettingsasclaude_effortand the claude block'soptionsaseffort, both absent when the flag was not used, because a run with the agent's own default is a different statement from a null. Docs: theclaudesection ofbench/harness/README.md.
Fixed
- The harness bench record's
placement.nodenow names the node that serves the model, not the master the CLI talked to.ainode-bench harnessbuilt the placement block by callingdescribe_via_httpwith the--ainodeweb base as both the control-plane and engine URL, so when DeepSeek V4 Flash was served on Spark-2 while the CLI pointed at Spark-1, the record'splacement.noderead "Spark-1". The CLI now resolves the serving node through the master's/api/server/statusloaded_models(the same fleet view the speed bench and the in-product bench use), callsdescribe_via_httpagainst that node's own web base, and records the engine port in placement. A model served on the master is unchanged; an unloaded model produces the same "does not report serving" warning. - The harness bench's OpenCode adapter now runs in the task's directory instead of the one the bench was started from, which is the whole of #118. Every
--harness opencodeinvocation from the runner exited 1 in about 1 s with{"type":"error",...,"message":"Unexpected server error. Check server logs for details."}on stdout and an empty log, while the identical command run by hand passed, so the adapter was carried as unverified and the DeepSeek harness record's OpenCode rows are marked not measured.--print-logs --log-level DEBUGoff a runner-shaped invocation named it in one run: opencode creates its instance in the working directory, reads the run's project-localopencode.jsonthere, and then creates a second instance and builds the session against the directory named byPWD- which a subprocess started withcwd=still inherits from its parent, so it was the shell the bench was launched in. That instance has noainode-benchprovider, so the session died withProviderModelNotFoundError: Model not found: ainode-bench/<model id>, reported on stdout as the generic server error. By hand it passed because a personcds in first andPWDis then correct; that asymmetry, not flakiness, is why it looked intermittent, and it is why the other four harnesses were never affected (none of them readsPWDto find its project). Two fixes, one cause: the argv now states--dir <workdir>, and_launchsetsPWDto the launch cwd and drops a staleOLDPWDfor every harness, since the wrong value also reaches whatever shell an agent spawns for its own tools. Verified by the runner itself,--only-tasks isogramagainst Ornith 1.5 35B A3B NVFP4 on the fleet: three consecutive runs, pass@1 1/1 at 25.5 s, 26.7 s and 26.5 s, zero crashes, where the same command on the previous commit crashed in 0.9 s twice out of two. - An OpenCode bench run no longer reads the operator's global config, and a failed one now says what failed.
OPENCODE_CONFIG_DIRoverrides the XDG config search outright, so the four isolatedXDG_*variables the adapter already set were bypassed by any exported value (Orca exports one) and somebody's global config, plugins and agents were inside the measurement; the overlay now points that variable at the run's own empty config directory too, the same ruleDSH_HOMEandCLAUDE_CONFIG_DIRfollow. The argv also carries--print-logs --log-level ERROR, which is diagnostics rather than behaviour - the logs go to stderr, stdout stays the NDJSON event stream the adapter parses - so a future failure reaches the record'sstderr_tailwith a named cause instead of the opaque "Unexpected server error" that #118 had to work from; verified against the old broken shape, where it prints theProviderModelNotFoundErrorin 1388 bytes. - The bind wait watches what the engine is DOING; log silence is only the second opinion (#112). vLLM prints nothing at all while it loads weight shards, runs torch.compile and autotunes FlashInfer, so "no log line for N seconds" could never tell a wedged engine from a working one: three healthy starts in two days were declared silent (Qwen3.8 27B NVFP4 quiet for 206 s between shard updates, Nemotron 3.5 Lightning for 363 s inside autotune and graph capture, then a 48-minute autotune), and the window had to grow 120 to 360 to 900 s to keep up. Since 0.5.13 that verdict broke nothing (the relaunch is a no-op while the container is alive) but it spent the engine's single retry and wrote a wrong reason into the log. Backends now publish
activity_mark(): forNvidiaBackendthe last time THAT engine's container was seen burning CPU, read from the container's cgroupcpu.statwhen the tree is visible and otherwise fromdocker stats --no-stream, cached so one bind poll costs one probe and run off the event loop thread. A log line OR an advancing activity mark resets the silence clock, and the start is called dead only when both have been quiet for the whole budget, which is why the budget comes back down to 300 s (engine_bind_log_silence_seconds): a truly wedged engine does no work at all and is caught in five minutes again instead of fifteen. The hard death signal (container exited) and the 1800 s ceiling are unchanged, and a backend that cannot see its engine's container reportsNoneand keeps the old log-driven behaviour. GPU utilization is deliberately not part of this: the pynvml path reads device 0 as a whole, so it cannot say which stacked engine is busy, and it reads 0% during exactly the phases that need covering. engine_readyin/api/nodesmeans the engine answers/v1/models, not that the node is reachable (the second half of #112). The field was derived fromstatus, which is heartbeat health (online/stale/offline) and true of every node whose orchestrator is up, becauseClusterNode.from_discoveredreplaces the announced engine state with that health: Spark-4 was publishedengine_ready: truewhile its engine was still loading and itsloaded_modelswas empty.ClusterNodenow carries the node's own word on its engine (engine_status, straight from the announcement, which that node's sync loop already sets from a live/v1/modelsprobe), the local row answers from a fresh localhost probe rather than thereadylatch, and a member node (which runs no local vLLM) still reads ready once discovered.statusis untouched, so the topology's online count keeps working.- The two ranks of an
mplaunch stay in step through FlashInfer autotune, instead of gloo killing the pair 48 minutes in (#134). Kernel warmup runs per rank. On Qwen3.8-Flash-Next the head, warm from an earlier launch, finished autotune in tens of minutes while the peer tuned 856 profiles from cold, sat 1800 s on its firsttrtllm::fused_moe::gemm1profile, and the barrier after autotune died withConnection closed by peeron one side andApplication timeout caused pair closureon the other. 1800 s is not a kernel hang: it is PyTorch's default timeout for the CPU (gloo) group, which is the group the barrier and the autotune sync both use. Three things change. The head now ships its compiled-kernel cache to a peer before that peer's container starts (_ensure_peer_has_jit_cache): one transfer per child of the recipe'sVLLM_CACHE_ROOT, skipped per child when the peer already has it, and a transfer that fails is a warning rather than a refused launch, since a missing cache costs minutes and a refused launch costs the model. The Flash-Next entry turns autotune off with--no-enable-flashinfer-autotune(the off form ofKernelConfig.enable_flashinfer_autotune, whichkernel_warmup.pyreads to skip the whole phase) and raises both collective floors to 5400 s, above the worst warmup measured on the pair; the stated cost is that its fused MoE and fp4 GEMMs now run FlashInfer's heuristic tactic rather than a measured one, so decode is slower than the 26.6 tok/s of the launch that got through. It also states the cache roots it was missing, so its JIT output persists per node inside the mounted HF cache instead of dying with the container and re-tuning from cold every launch. The bind wait's ceiling goes from 1800 s to 3600 s, because a launch that took 36 minutes to ready was being killed by the ceiling as well; the ceiling is the last resort and the activity signal from #112 still calls a genuinely wedged engine dead inside five minutes, so nothing waits an hour on a dead start. Worth writing down for the next person: seeding a peer's autotune file cannot help on its own. Only rank 0 reads that file and broadcasts its bytes to the other ranks, and the path holds a sha256 of the wholeVllmConfighash, so a hand-copied directory is both unread and easy to orphan. DeepSeek V4 Flash keeps its own autotune and timeouts exactly as proven. - Catalog entries state their active parameter count and architecture, so a bench record stops reporting a MoE as dense.
ModelInfohad no field for either, and the record'smodelblock was filled by reading theA<n>Bmarker off the model id, which works forOrnith-1.5-35B-A3Band not at all for an id that does not carry one: DeepSeek V4 Flash (284B total, 13B active) and Qwen3.8-Flash-Next (125B, 6B) were both written down asparams_b 125/284, active_b 125/284, arch dense, which on GB10 is the one number that predicts decode speed, said wrong.ModelInfonow carriesactive_params_bandarch("moe"/"dense"), every curated and fallback entry states its shape, and the model block takes both from the catalog when it has them and falls back to the id only for a model the catalog does not describe. An entry that says MoE without an active count (GLM-5.2 REAP, where pruning moves it) leavesactive_bout and both renderers print504B MoErather than claiming dense. The catalog cache round-trips the new fields and a cache file written before they existed still loads. Applies to the throughput bench, the in-product/api/benchruns and the harness bench alike: all three build the block throughainode/bench/fleet.py.
Image: ghcr.io/getainode/ainode:0.5.20