Skip to content

AINode 0.5.18

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 16 Sep 16:46
· 85 commits to main since this release
839fbdf

The distributed launch serves a model downloaded through AINode; harness bench; Qwen3.8-Flash-Next entry.

Fixed

  • A distributed launch now serves a model that was downloaded through AINode instead of making every node re-download it. POST /api/models/download-repo writes a flat <models_dir>/<owner--name> directory, not the Hugging Face cache layout, and the solo launch has served that directory for releases. The distributed shapes did not: both served the repo id, the head and peer containers mounted only the HF cache, and _ensure_peer_has_model distributed only a hub/models--<owner>--<name> entry, so a 133 GB model already sitting on the head was pulled again on every rank, over the WAN, for every launch. Both shapes now go through one resolver: when that flat directory exists on the head and the mount is trustworthy (the same AINODE_HOST_HOME test the solo path uses), every rank serves /ainode-models/<owner--name> with --served-model-name <repo id> so /v1/models is unchanged, and every rank mounts its own copy of the store there: the head's models_dir, a peer's /home/<ssh_user>/ainode-nvidia-models (chosen like the peer HF cache, because we ssh in as that user). _ensure_peer_has_model ships the flat directory over the fabric (rsync when present, else tar over ssh, skipped when the peer already has it) instead of the hub entry, and the peer's mkdir -p covers the new path so docker can never invent a root-owned empty bind source. With no such local copy, nothing changes: the repo id is served and only the HF cache is mounted. A recipe that states --served-model-name itself still wins, in solo too.

Added

  • Catalog entry: Qwen3.8-Flash-Next (NVFP4), nvidia/Qwen3.8-Flash-Next-NVFP4, as a two-node mp launch. Frontier MoE, 125B total and 6B active per token, plus a 51B PLE n-gram embedding and a 4B MTP module; 133 GB of mixed-precision weights (NVFP4 routed experts, FP8 elsewhere, modelopt), so it does not fit one 121 GB GB10 node. The entry pins vllm/vllm-openai:v0.29.0: the architecture and the FP8-PLE loader for mixed ModelOpt checkpoints landed in 0.29, and the fleet's 0.27.1 default and the local 0.28.0 images do not know it. Carries the card's recipe minus what the backend emits itself (--quantization modelopt, --enable-prefix-caching, --reasoning-parser qwen3, --tool-call-parser qwen3_coder, --enable-auto-tool-choice), fp8 KV cache, 262K context, --trust-remote-code, gmu 0.85. MTP speculative decoding is deliberately not enabled: it wants --enable-expert-parallel, which hangs on this MoE and hardware, so it is a follow-up. verified=false and recommended=false until it is served end to end on the pair. It is the strongest coding model in the Qwen3.8 line, ahead of the 27B and DeepSeek V4 Flash on three of the four coding rows in Qwen's own table.
  • A harness bench: scripts/ainode-bench.py harness measures whether a served model can drive a coding agent to passing tests, as opposed to how fast it generates. Ten Python practice exercises vendored from exercism/python (MIT, the same source Aider's polyglot benchmark draws on) are handed to a real agent CLI, and the score is what pytest says afterwards: pass@1, pass@2, mean wall seconds and crash count per harness. Adapters ship for aider, dsh, pi and opencode; every one is pointed at an AINode endpoint (http://<node>:3000/v1) and needs no API key beyond a placeholder. The hidden tests are copied into the working directory only after the harness has exited and removed again before the next attempt, so a harness can never read the assertions it is being judged on, and the second attempt is given the real test output the way Aider's protocol does it. --dry-run prints the exact argv, environment and config for every task and harness without running anything, and is the only part that works in a sandboxed shell: the agents themselves need normal network and home-directory access, and a sandboxed run hangs rather than failing loudly. dsh gets its own DSH_HOME because it validates every configured provider route at boot, so one stale entry in a person's ~/.dsh/settings.yaml would end every run with a transport error whatever provider the run selected. Records are schema 1 with a new top-level harness block and no results block, so scripts/render-bench-table.py skips them rather than rendering a coding run as a very slow model. Docs: bench/harness/README.md.

Image: ghcr.io/getainode/ainode:0.5.18