Repository navigation
score
score() rates instead of choosing. You describe a spectrum as an ordered list of levels, hand it a state, and it comes back with where the state lands on that line: one number that can fall between two of your levels, the probability of every level and how sure the model was.
It runs on the same single forward pass as choice and is exactly as cheap; the answer arrives as a digit instead of a letter. The mechanics are in Under the hood.
Use score() when the answer is a position on a spectrum that makes sense as distinct, describable steps: how fast a reply is needed, how warm a lead is, how close a draft is to done.
If the answer is instead one of a fixed set with no order between them, use choice. If it is a single yes/no proposition, use noul.
criteria is an array of level descriptions ordered from the low end of the spectrum to the high end. A level's number is its position in the array, starting at 0. Between 2 and 10 levels are supported.
Writing levels well is most of using score() well, and the one rule that matters is this: describe situations, not degrees. Each level is what the model weighs the state against on its own, so a description that only makes sense next to another level ("worse than the previous one", "twice as urgent") describes nothing. Concrete situations give the model something to picture:
// Good: each level is a situation you could paste into a book
criteria: [
"No rush; whenever there is a spare moment",
"Within the week is fine",
"Before the day ends",
"Within the hour",
"Right now, drop everything",
];// Weak: names for degrees rather than situations
criteria: [
"Not very urgent",
"Somewhat urgent",
"Very urgent",
"Extremely urgent",
];Two additional habits that show up in real runs:
- Neighbor levels that overlap split the answer between them. You will see it in the answer as mass shared between two adjacent numbers; if two levels would happily both apply, one of the descriptions is too wide.
- Cover the ends of the spectrum honestly. A state the levels never quite match is exactly the case that comes back with no signal (see below), and that is what you want instead of a forced fit.
import { score } from "smart-decisions";
const model = {
apiBaseUrl: "http://localhost:8000/v1",
apiKey: "a-super-secret-api-key",
model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
};
const answer = await score({
model,
state:
"Can you hop on a quick call before the 3pm? Legal is asking about the rider we flagged this morning.",
instructions: "How fast does this need a reply?",
criteria: [
"No rush; whenever there is a spare moment",
"Within the week is fine",
"Before the day ends",
"Within the hour",
"Right now, drop everything",
],
});
console.log(answer);The answer:
{
score: 2.43; // position on the level line
probabilities: { '0': 0.0, '1': 0.0, '2': 0.57, '3': 0.43, '4': 0.0 }; // sums to 1
confidence: 0.38; // 0..1 — flat distribution → low, single peak → high
legend: {
'0': 'No rush; whenever there is a spare moment',
'1': 'Within the week is fine',
'2': 'Before the day ends',
'3': 'Within the hour',
'4': 'Right now, drop everything',
};
}-
scoreis the position on the level line: each level number weighted by its probability, added up, the probability weighted mean of the level numbers. That is why 2.43 is a legal answer even though 2.43 is not a level: most of the mass sits on level 2 with the rest bleeding into level 3. The same arithmetic also means a score of 1.0 can be all probability on level 1, or half on each of levels 0 and 2. When a decision rests on it, readprobabilitiesalongside the score instead of rounding:Math.round()can quietly land on a level nobody voted for. When you need one concrete level, take the highest entry fromprobabilities; when you need a place to put thresholds, the score is the better driver. -
probabilitiescarries one probability per level, keyed by level number as a string ("0", "1", ...), and sums to one. -
confidenceis the same measure as onchoice(): near 0 on a flat distribution, near 1 when one level holds the mass. A middling confidence is not automatically a failure here; a state sitting honestly between two levels produces exactly that. Gate autonomous action on it the same way you would on a choice. -
legendmaps each level number back to its description. It exists for the humans reading your logs or your UI, who need the text next to the number.
The optional fields (mode, maxRetries, timeoutMs, model.extraBody) mean the same thing as on choice, which details them — and mode selects between the same three passes: 'system1', 'system2' (Under the hood) and 'auto' (Auto mode). A request in practice:
import { score } from "smart-decisions";
const model = {
apiBaseUrl: "http://localhost:8000/v1",
apiKey: "a-super-secret-api-key",
model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
extraBody: { reasoning_effort: "none" }, // sent verbatim; works where the backend understands it
};
const answer = await score({
model,
state:
"Lead replied to the pricing email ninety minutes after it was sent, asking for a demo with the full engineering team for a rollout next quarter.",
instructions: "How close is this lead to buying?",
criteria: [
"A tire kicker with no follow through",
"Curious, but no timeline commitment",
"Actively evaluating, asking scoped questions",
"Talking rollout and timelines, wants the team involved",
],
mode: "system1",
maxRetries: 4,
timeoutMs: 30_000,
});On a five level spectrum a readout like 2.8 says this lead sits between "actively evaluating" and "talking rollout", leaning upwards. That is a signal for your CRM logic to act on, and the confidence tells you whether to route the lead to a person at all.
- Fewer than 2 or more than 10 levels throw.
- A level that is not a string throws.
- In System 2 mode, stays failing structural replies past the retry budget throw with the last rejection reason.
choice for unordered option sets, noul for yes/no judgments, Auto mode for testing when the answer is worth deliberating on, and Under the hood for how one token becomes a distribution.
Documentation
Examples
Reference