Repository navigation
AINode 0.5.32
Decisions use the model's own calibration: /v1/decide and /v1/systemone apply a decision adapter's fitted temperatures and say so in every answer, a decision model warms its answer grammar the moment it binds, the decision bench scores any model on the public Jevals sets with the boards' own formulas, and a master that serves no model is safe to run, which is the first step of moving the fleet's control plane to Atlas.
Added
/v1/decideand/v1/systemoneapply a decision adapter's own temperatures (#276, #281). When the served model's directory carriestemperatures.json, each question's label logprobs are divided by the temperature fitted for its kind (choice,noul,score) before the softmax. Both responses reportcalibration: {applied, temperatures}, and"calibration": "raw"returns the engine's own spread.- A decision model warms its answer grammar as soon as its engine binds (#277, #281). One small constrained question per kind is sent in the background, with each compile time logged.
/api/statusgainsinstances[]withwarmandwarm_compile_secondsper instance. - The decision bench scores models on the public question sets the independent Jevals boards use, with the boards' own formulas (#272), so an AINode-served model can be compared with Jev and its clones on one recipe.
ainode-bench decide --suite all --transport decide|systemoneruns PubMedQA (noul, 300), Banking77 (choice, 300, K=77), HelpSteer2 helpfulness (score, 300) andLocalLLaMA/typed-decisions(mixed, 2000 decisions over 400 cases) at 5 repeats per question, against either AINode's ownPOST /v1/decideor any endpoint speaking the JevPOST /v1/systemonewire format. Per set it reports the Decision Score against the label prior (100 perfect, 0 the base rates, negative worse), accuracy with its guessing floor, the calibration gap in points over ten bins with the reliability table behind it, the hand-off share at 95 percent, the published gate with this run's coverage, pick flips and confidence swing across the repeats, p50 and p95 latency, questions per second, malformed answers, and cost when a price is given.--questions <file>scores any question file in the same shape, so a private blind set never has to be committed. The recipe, the date each source page was read and every deviation from it are inbench/decide/JEVALS.md; the sets arebench/decide/sets/andainode-bench decide downloadfetches the item text, which is not committed, and verifies every state against its published hash.
Changed
- A decision request that meets a cold grammar compile gets a 503 saying so, and has longer to finish (#281). The engine-call limit is 300 s, up from 180, and a call that times out after connecting reads "the engine is compiling the answer grammar, retry" instead of "unreachable".
ainode doctorpasses on a node that loads no model (#282). "No model loaded" is INFO, and docker and engine-backend findings are INFO on such a node, so a routing-only node whose GPU belongs to something else no longer fails. A node that pins a model still fails exactly as before.
Fixed
- A request with no
modelon a node that serves no model of its own gets a 400missing_model_fieldnaming the field (#282), instead of being routed to a model the node does not serve. Nodes that serve a primary keep their default. - Unloading or ejecting a node's primary no longer promotes another instance (#282), for example Whisper, into
config.modelwith the old primary's engine parameters. The primary slot and its overrides are cleared, and stacked instances stay on their own ports. - An empty account list is never replicated (#282). The master's push, a worker's pull, the CLI push and
/api/auth/users/sync(409empty_account_list) all refuse it and log how to fix the store, so a master with an emptyusers.jsoncan no longer sign every dashboard user out of the fleet.
Image: ghcr.io/getainode/ainode:0.5.32