Repository navigation
AINode 0.5.9
Distributed launch without Ray, interface autodetect, and the DeepSeek V4 Flash recipe.
Added
- A distributed launch that needs no Ray in the engine image: the vLLM
mp
multi-node shape (#84). The distributed path could only do one shape, a
ray start --headcontainer plus SSH-launchedray startworkers plus a
docker execofvllm serve --distributed-executor-backend rayinside the
head, so it required therayCLI in the engine image. Two engine images we
actually need do not have it: stockvllm/vllm-openai:v0.27.1(the head
container exits 127, "ray: command not found") and the custom GB10 build that
is the only thing serving DeepSeek V4 Flash correctly on sm121.NodeConfig
gainsdistributed_executor(rayby default, ormp), and thempshape
runs onevllm servecontainer per node: rank 0 on the head,--node-rank k --headlesson each peer over SSH, all rendezvousing on
--master-addr/--master-portwith vLLM's own multi-node executor. Peers go
out before the head because the rendezvous, not a Ray port, is what waits.
Container shape is the proven one:--network host --ipc host --shm-size 64g --ulimit memlock=-1 --ulimit stack=67108864 --gpus all, plus
--device /dev/infinibandwhen the host has it. Serve args come from the same
builder as every other launch, so a recipe flag still suppresses the built-in
rather than duplicating it. Readiness,last_log_activity, the adaptive bind
wait andstop()all work with nodocker exec'd process: the head container
is the server, itsdocker logs -fis the engine log, and teardown removes the
head plus every peer container.NodeConfig.extra_volumesis new alongside it
(extrahost:container[:ro]mounts, applied to the solo and distributed
commands) for an engine image that wants a writable cache outside the HF cache. - The distributed launch is recipe-aware (#84).
POST /api/sharding/launch
built its per-instance config from the sharedNodeConfigand ignored the
catalog, so a curated model launched across nodes on the fleet default image
with none of its proven flags, while the same model loaded on one node got its
whole recipe. It now resolves the catalog recipe (engine image, extra vLLM
args, env, volumes, distributed shape, KV-cache dtype, context length,
trust-remote-code, recommended GPU fraction) and applies it as defaults, with
the same per-launch body keys the solo load accepts overriding it, plus
distributed_executorfor the shape. Body parsing is now one shared
validator, so both paths reject a malformed recipe with the same 400 instead of
silently dropping it. The instance record carriesdistributed_executorand a
primary launch persists the shape toconfig.json, so a restart replays the
same shape rather than falling back to Ray on the default image. - Catalog entry: DeepSeek V4 Flash (DSpark, FP8) (#84), id
deepseek-v4-flash-dspark,fraserprice/DeepSeek-V4-Flash-DSpark. Frontier
MoE, 284B total with 13B active per token, 1M context, MIT,proven_tp=2. It
carries the full two-node recipe:distributed_executor="mp", the GB10 engine
image,nvfp4_ds_mlaKV cache, DSpark speculative decoding, the deepseek_v4
tokenizer/tool-call/reasoning parsers, and the engine env the image needs
(its ENTRYPOINT is empty,HOMEis/tmpand vllm lives at
/opt/env/bin/vllm, soPATH, the CUDA paths,HF_HOMEand the JIT cache
dirs are all stated, the last pointed inside the HF cache mount so compiled
kernels persist per node without a fleet-specific path).verified=False
until the two-node serve is proven live, and it needs that GB10 vLLM build
present on every node: it is a local image today, a registry publish is a
follow-up.
Fixed
VLLM_ATTENTION_BACKEND=TRITON_ATTNis no longer forced onto a custom or
newer engine image (#84). It was injected into every engine container. On the
pinned 0.17 default it is a documented no-op hedge, and vLLM 0.27/0.28 merely
log it as an unknown variable, but the 0.21-based GB10 fork that serves
DeepSeek V4 Flash does honor it, and pinning a dense attention backend over
that model's sparse MLA path is exactly the kind of override that makes a serve
produce confident nonsense. The pin now sits behind the same gate as the other
0.17-era workarounds (_is_pinned_default_image), so behaviour on the default
image is byte-identical, including the systemd env override, and every other
image gets no attention override unless its recipe'sextra_envstates one.- The cluster interface is auto-detected instead of guessed (#34, #61). The
installer wrote the DGX Spark NIC nameenP2p1s0f1np1into every new
config.jsonandNodeConfig.cluster_interfacedefaulted toeno1, so on an
ASUS GX10 or any other box neither name existed. NCCL, Ray, Gloo and UCX were
pointed at a device that was not there: fabric-IP detection returned nothing
and the engine either bound127.0.0.1silently or failed EngineCore init
with no usable reason. The reporter on #34 had to find the real name
(enp1s0f0np0) withip -br addrand hand-editconfig.json. AINode now
ranks this host's real interfaces, preferring an up RDMA-capable port with an
IPv4, then the default-route device, then any up non-virtual device, and
skipping loopback, bridges, veth pairs and VPN or overlay tunnels. The
configured name still wins whenever it exists, so a pinned interface is never
overruled; when it does not exist AINode logs one warning naming both the
configured and the chosen device.cluster_interfacenow defaults to empty,
meaning auto-detect, the installer detects a name at install time with the
same ranking,ainode startprints the chosen interface and its address on a
Fabricline, and the "could not detect fabric IP" error now lists the
interfaces that do have an address. ainode starton a host with no vLLM fails with an explanation, not a
traceback (#61). A pip-installed AINode outside the container runs the eugr
backend by default, which shells out tovllm serve, so the start died with a
rawFileNotFoundError: [Errno 2] No such file or directory: 'vllm'. The
start now checks for vLLM first and exits 1 with the install command and the
engine_backend: nvidiaalternative, the backend turns thatPopenfailure
into a clear error carrying the same guidance, and the old "using the eugr
container backend" line is gone: eugr is the in-container path, not a way to
run a container from the host.- Replay no longer launches every engine twice (#80). Engines run with
--rm, so afterdocker stopthe daemon removes them asynchronously and
docker rm -freturns first; adocker run --namein that gap failed with
"Conflict. The container name ... is already in use". On the 0.5.8 roll every
engine's first launch died that way at 0 s and the adaptive wait relaunched
it. The pre-launch cleanup and the replay's orphan sweep now poll until the
name is actually gone (bounded at 90 s, logged if it never clears).
Image: ghcr.io/getainode/ainode:0.5.9