Skip to content
expilu edited this page Oct 3, 2026 · 3 revisions

score()

score() rates instead of choosing. You describe a spectrum as an ordered list of levels, hand it a state, and it comes back with where the state lands on that line: one number that can fall between two of your levels, the probability of every level and how sure the model was.

It runs on the same single forward pass as choice and is exactly as cheap; the answer arrives as a digit instead of a letter. The mechanics are in Under the hood.

When to use it

Use score() when the answer is a position on a spectrum that makes sense as distinct, describable steps: how fast a reply is needed, how warm a lead is, how close a draft is to done.

If the answer is instead one of a fixed set with no order between them, use choice. If it is a single yes/no proposition, use noul.

Levels: the shape of the spectrum

criteria is an array of level descriptions ordered from the low end of the spectrum to the high end. A level's number is its position in the array, starting at 0. Between 2 and 10 levels are supported.

Writing levels well is most of using score() well, and the one rule that matters is this: describe situations, not degrees. Each level is what the model weighs the state against on its own, so a description that only makes sense next to another level ("worse than the previous one", "twice as urgent") describes nothing. Concrete situations give the model something to picture:

// Good: each level is a situation you could paste into a book
criteria: [
  "No rush; whenever there is a spare moment",
  "Within the week is fine",
  "Before the day ends",
  "Within the hour",
  "Right now, drop everything",
];
// Weak: names for degrees rather than situations
criteria: [
  "Not very urgent",
  "Somewhat urgent",
  "Very urgent",
  "Extremely urgent",
];

Two additional habits that show up in real runs:

  • Neighbor levels that overlap split the answer between them. You will see it in the answer as mass shared between two adjacent numbers; if two levels would happily both apply, one of the descriptions is too wide.
  • Cover the ends of the spectrum honestly. A state the levels never quite match is exactly the case that comes back with no signal (see below), and that is what you want instead of a forced fit.

A full walkthrough

import { score } from "smart-decisions";

const model = {
  apiBaseUrl: "http://localhost:8000/v1",
  apiKey: "a-super-secret-api-key",
  model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
};

const answer = await score({
  model,
  state:
    "Can you hop on a quick call before the 3pm? Legal is asking about the rider we flagged this morning.",
  instructions: "How fast does this need a reply?",
  criteria: [
    "No rush; whenever there is a spare moment",
    "Within the week is fine",
    "Before the day ends",
    "Within the hour",
    "Right now, drop everything",
  ],
});

console.log(answer);

The answer:

{
  score: 2.43; // position on the level line
  probabilities: { '0': 0.0, '1': 0.0, '2': 0.57, '3': 0.43, '4': 0.0 }; // sums to 1
  confidence: 0.38; // 0..1 — flat distribution → low, single peak → high
  legend: {
    '0': 'No rush; whenever there is a spare moment',
    '1': 'Within the week is fine',
    '2': 'Before the day ends',
    '3': 'Within the hour',
    '4': 'Right now, drop everything',
  };
}
  • score is the position on the level line: each level number weighted by its probability, added up, the probability weighted mean of the level numbers. That is why 2.43 is a legal answer even though 2.43 is not a level: most of the mass sits on level 2 with the rest bleeding into level 3. The same arithmetic also means a score of 1.0 can be all probability on level 1, or half on each of levels 0 and 2. When a decision rests on it, read probabilities alongside the score instead of rounding: Math.round() can quietly land on a level nobody voted for. When you need one concrete level, take the highest entry from probabilities; when you need a place to put thresholds, the score is the better driver.
  • probabilities carries one probability per level, keyed by level number as a string ("0", "1", ...), and sums to one.
  • confidence is the same measure as on choice(): near 0 on a flat distribution, near 1 when one level holds the mass. A middling confidence is not automatically a failure here; a state sitting honestly between two levels produces exactly that. Gate autonomous action on it the same way you would on a choice.
  • legend maps each level number back to its description. It exists for the humans reading your logs or your UI, who need the text next to the number.

An example with the optional parameters

The optional fields (mode, maxRetries, timeoutMs, model.extraBody) mean the same thing as on choice, which details them — and mode selects between the same three passes: 'system1', 'system2' (Under the hood) and 'auto' (Auto mode). A request in practice:

import { score } from "smart-decisions";

const model = {
  apiBaseUrl: "http://localhost:8000/v1",
  apiKey: "a-super-secret-api-key",
  model: "/models/Qwen3.5-4B-Q4_K_M.gguf",
  extraBody: { reasoning_effort: "none" }, // sent verbatim; works where the backend understands it
};

const answer = await score({
  model,
  state:
    "Lead replied to the pricing email ninety minutes after it was sent, asking for a demo with the full engineering team for a rollout next quarter.",
  instructions: "How close is this lead to buying?",
  criteria: [
    "A tire kicker with no follow through",
    "Curious, but no timeline commitment",
    "Actively evaluating, asking scoped questions",
    "Talking rollout and timelines, wants the team involved",
  ],
  mode: "system1",
  maxRetries: 4,
  timeoutMs: 30_000,
});

On a five level spectrum a readout like 2.8 says this lead sits between "actively evaluating" and "talking rollout", leaning upwards. That is a signal for your CRM logic to act on, and the confidence tells you whether to route the lead to a person at all.

Error cases worth knowing

  • Fewer than 2 or more than 10 levels throw.
  • A level that is not a string throws.
  • In System 2 mode, stays failing structural replies past the retry budget throw with the last rejection reason.

See also

choice for unordered option sets, noul for yes/no judgments, Auto mode for testing when the answer is worth deliberating on, and Under the hood for how one token becomes a distribution.

Clone this wiki locally