Repository navigation
choice
choice() is the main feature of the library. It answers a question of the form "given this situation and these options, which one is it now?": you hand it a state, a fixed set of options and some instructions, and it comes back with the winning option, the probability of every option and how sure the model was about the pick.
It runs on any LLM behind an OpenAI-compatible v1 API, hosted or self hosted. The only real requirement is logprobs support: the model generates exactly one token and the probability distribution behind that single token is the answer. One decision costs one forward pass, which is why it lands in the milliseconds on modest hardware. The mechanics are in Under the hood.
Use choice() when the answer is one option out of a set with no order between them: the department for a support ticket, the language for a reply, the next action for an agent. Your code owns the set; the model can only pick inside it.
If the answer is instead a position on a spectrum you can describe in distinct steps (how urgent, how warm, how close to done), use score. If the question is a single yes/no proposition, use noul.
import { choice } from "smart-decisions";
const model = {
apiBaseUrl: "http://localhost:8000/v1", // your API url. i.e. your llama.cpp server
apiKey: "a-super-secret-api-key", // as required or not by your provider
model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
};
const answer = await choice({
model,
state: "It is raining and I am at home. I'm bored.",
instructions: "Give me a good plan to do now",
criteria: {
walk: "Go for a walk",
movie: "Watch a movie",
beach: "Go to the beach",
},
});
console.log(answer);Three fields carry the question:
-
stateis the situation. Whatever context the decision depends on belongs in here. -
instructionsis what you want the model to do with that state. One clear sentence works best. -
criteriais the option set. The keys are names your code will read back; the values are the descriptions the model judges each option by.
The call returns:
{
choice: 'movie'; // option name with the highest probability
probabilities: { walk: 0.00081, movie: 0.99913, beach: 0.00006 }; // sums to 1
confidence: 0.9935; // 0..1 — flat distribution → low, single peak → high
}-
choiceis the option name with the highest probability: the key itself, from your owncriteria, never a letter or an index. -
probabilitiesis the full distribution over the options and always sums to one. Read the shape, not just the champion: how much mass the runner up carries tells you how contested the pick was. -
confidencesummarizes the same distribution: it stays near 0 when the probability spreads evenly (the model has no clear instinct) and climbs towards 1 as the mass piles up on one option. This is the number your software gates on: low confidence is the signal to escalate to a human, system2 or another code path instead of acting onanswer.choicealone.
Required:
-
model: the model to query and the OpenAI-compatible v1 API serving it. It carriesapiBaseUrl,apiKey,modeland, optionally,extraBody. One object can back any number of questions, so define it once. -
state: the situation to evaluate, as a string. -
instructions: the question to answer, as a string. -
criteria: the options, as{ optionName: description }. Between 2 and 26 options: each gets one letter of the alphabet as its answer symbol.
Optional:
-
mode:'system1'(the default; the one token pass described Under the hood),'system2'(the deliberated structured answer: Under the hood, which is System 2's section down the page), or'auto'(fast first, deliberate only when unsure: Auto mode). -
thinking: turn the reasoning trace of the System 2 answer on or off (boolean). Ignored in System 1 — a one token answer cannot reason, and the type system stops it from being set there at all. Explained under the hood above. -
autoModeThreshold: the escalation threshold formode: 'auto'. Defaults to 0.7. Auto mode explains reading it. -
debug: log the internals of the answer to stderr (Under the hood). Works on both modes. -
maxRetries: how many times a failed attempt (network errors and 408, 409, 429, 5xx responses) is retried, with exponential backoff. Defaults to 2. -
timeoutMs: per attempt timeout, in milliseconds. Defaults to 600000, ten minutes. -
model.extraBody: fields forwarded verbatim into the request body for engine or model specific settings the library does not model. Showcased in the example below.
Beyond the option count, inputs are validated at runtime because plain JavaScript callers can pass anything: criteria values must be strings, and no option set gets silently repaired.
import { choice } from "smart-decisions";
const answer = await choice({
model: {
apiBaseUrl: "http://localhost:8000/v1",
apiKey: "a-super-secret-api-key",
model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
// extraBody is forwarded verbatim into the request body. Only the keys your
// backend understands take effect; lenient engines ignore the rest.
extraBody: { reasoning_effort: "none" },
},
mode: "system1", // the default, the fast one token pass
maxRetries: 0, // retry nothing: interactive paths prefer failing fast (default 2)
timeoutMs: 10_000, // per attempt, milliseconds (default 600000)
state: "It is raining and I am at home. I'm bored.",
instructions: "Give me a good plan to do now",
criteria: {
walk: "Go for a walk",
movie: "Watch a movie",
beach: "Go to the beach",
},
});A word on extraBody, since it is the one field where every backend differs. The library already disables thinking by default (chat_template_kwargs: { enable_thinking: false }), which is what you want on llama.cpp, vLLM and SGLang for a one token answer. Other engines expose the same idea under different keys: reasoning_effort: 'none' on OpenAI, OpenRouter and Ollama, think: false on Ollama's native API. Two things to know: keys the library itself depends on (model, messages, stream, logprobs, top_logprobs, max_tokens, temperature) cannot be overridden through extraBody, and engines that validate the request body strictly (the hosted OpenAI family) reject unknown fields with HTTP 400 instead of ignoring them.
mode: 'system2' runs the deliberate pass: the model reasons and rates every option 0..10 in a minimal structured reply, and the ratings become the distribution, same ChoiceAnswer, same entropy confidence:
const answer = await choice({
...,
mode: "system2",
thinking: false, // true for activating reasoning on the answer (and more time to answer...)
});- Fewer than 2 options throws: there is nothing to choose between.
- More than 26 options throws in System 1 mode: one letter per option is the hard ceiling.
- A criteria value that is not a string throws. The runtime check exists because plain JavaScript callers can pass anything.
- An API response with no logprobs at all throws rather than inventing a distribution.
- In System 2 mode, a reply that keeps failing schema validation past the retry budget throws with the last rejection reason.
score, noul, Under the hood, Auto mode, and the worked example in Support ticket routing.
Documentation
Examples
Reference