Skip to content

AINode 0.5.32

Choose a tag to compare

@webdevtodayjason webdevtodayjason released this 26 Sep 09:30
· 6 commits to main since this release
ae87d20

Decisions use the model's own calibration: /v1/decide and /v1/systemone apply a decision adapter's fitted temperatures and say so in every answer, a decision model warms its answer grammar the moment it binds, the decision bench scores any model on the public Jevals sets with the boards' own formulas, and a master that serves no model is safe to run, which is the first step of moving the fleet's control plane to Atlas.

Added

  • /v1/decide and /v1/systemone apply a decision adapter's own temperatures (#276, #281). When the served model's directory carries temperatures.json, each question's label logprobs are divided by the temperature fitted for its kind (choice, noul, score) before the softmax. Both responses report calibration: {applied, temperatures}, and "calibration": "raw" returns the engine's own spread.
  • A decision model warms its answer grammar as soon as its engine binds (#277, #281). One small constrained question per kind is sent in the background, with each compile time logged. /api/status gains instances[] with warm and warm_compile_seconds per instance.
  • The decision bench scores models on the public question sets the independent Jevals boards use, with the boards' own formulas (#272), so an AINode-served model can be compared with Jev and its clones on one recipe. ainode-bench decide --suite all --transport decide|systemone runs PubMedQA (noul, 300), Banking77 (choice, 300, K=77), HelpSteer2 helpfulness (score, 300) and LocalLLaMA/typed-decisions (mixed, 2000 decisions over 400 cases) at 5 repeats per question, against either AINode's own POST /v1/decide or any endpoint speaking the Jev POST /v1/systemone wire format. Per set it reports the Decision Score against the label prior (100 perfect, 0 the base rates, negative worse), accuracy with its guessing floor, the calibration gap in points over ten bins with the reliability table behind it, the hand-off share at 95 percent, the published gate with this run's coverage, pick flips and confidence swing across the repeats, p50 and p95 latency, questions per second, malformed answers, and cost when a price is given. --questions <file> scores any question file in the same shape, so a private blind set never has to be committed. The recipe, the date each source page was read and every deviation from it are in bench/decide/JEVALS.md; the sets are bench/decide/sets/ and ainode-bench decide download fetches the item text, which is not committed, and verifies every state against its published hash.

Changed

  • A decision request that meets a cold grammar compile gets a 503 saying so, and has longer to finish (#281). The engine-call limit is 300 s, up from 180, and a call that times out after connecting reads "the engine is compiling the answer grammar, retry" instead of "unreachable".
  • ainode doctor passes on a node that loads no model (#282). "No model loaded" is INFO, and docker and engine-backend findings are INFO on such a node, so a routing-only node whose GPU belongs to something else no longer fails. A node that pins a model still fails exactly as before.

Fixed

  • A request with no model on a node that serves no model of its own gets a 400 missing_model_field naming the field (#282), instead of being routed to a model the node does not serve. Nodes that serve a primary keep their default.
  • Unloading or ejecting a node's primary no longer promotes another instance (#282), for example Whisper, into config.model with the old primary's engine parameters. The primary slot and its overrides are cleared, and stacked instances stay on their own ports.
  • An empty account list is never replicated (#282). The master's push, a worker's pull, the CLI push and /api/auth/users/sync (409 empty_account_list) all refuse it and log how to fix the store, so a master with an empty users.json can no longer sign every dashboard user out of the fleet.

Image: ghcr.io/getainode/ainode:0.5.32