This release turns the provider from something that declares models into something that can run them, and adds the machinery to decide which ones are worth running.
Run model servers from the CLI
models serve, models stop and models ps start and stop the backends a node serves from. Previously setup printed a docker command for you to run by hand.
computing-provider models serve --backend llamacpp \
--weights /models/Qwen3.8-27B-UD-Q4_K_S.gguf --gpus 2,3 \
--context-length 65536 --port 30001 \
Qwen/Qwen3.8-27B --parallel 2 -fa on -ctk q8_0Backends: llamacpp, vllm, sglang. Ownership is decided by container labels alone, so a backend you started yourself is invisible to ps and cannot be stopped by this tool. Three refusals protect an existing setup, all before a container is created: a managed container already serving the model, a models.json entry with no managed container behind it, and a name collision with a container this tool did not create. Servers publish on 127.0.0.1 only — a model server has no authentication of its own. --dry-run prints the docker command without running it.
Know what a model will cost before loading it
The marketplace publishes no VRAM requirement for any model, so the node derives one. Three tiers, descending trust: measured (it has run here), derived (computed from published metadata), unknown (refused). A model that cannot be sized is never selected automatically.
Getting this right needed corrections a generic formula does not make — hybrid attention, multi-head latent attention, sliding windows, and measured quantisation sizes. Validated against a real node: derived 17.43 GiB against 17.42 measured.
models serve records what a model actually used, in model-memory.json, and regenerates cp.md from it — including the exact command that produced each record.
Plan and automate model switching
inference plan runs one cycle of the auto-switch decision and prints what would happen, changing nothing. A local model proposes; ordinary code decides whether the proposal is allowed. Guardrails refuse a whole plan rather than half of one.
[Inference.AutoSwitch] Enable = true starts the scheduler that acts on it, with dwell time, agreement across cycles, and a daily switch cap. Off by default, and it refuses to start rather than starting degraded.
Ask a local model about the node
computing-provider agent "<goal>" gives a goal in plain language to a model this node already serves, with a fixed set of eight tools. Bounded by its tools, not by its judgement: no shell, no way to reach anything outside the registry. Tools that change the node are refused without --allow-actions, are refused entirely when the run is not attached to a terminal, and still hit every guardrail a person would.
Finishing with zero successful tool results is refused in code — an early build invented a plausible earnings table without calling a single tool.
Alerts on every switch
Whenever the set of served models changes, you are mailed through your configured [Alerts] transports, carrying who decided and why.
Fixes
--gpus 2,3failed outright withcannot set both Count and DeviceIDs; docker CSV-parses the value, so any multi-GPU serve was brokenconfig.toml's model list is kept in step withmodels.json, which the self-check audit compares- a re-enabled model reported
state: disabledfor the life of the process while serving normally - an unpublished VRAM requirement is treated as unknown rather than as a fit — the previous rule passed 600B models onto 10 GB cards
- dashboard: finished earnings history is no longer re-priced from live counters, and the earnings breakdown shows one line per model
Upgrading
No configuration changes are required. The scheduler and the agent are both opt-in.
Full Changelog: v0.5.5...v0.6.0