Skip to content

Auto mode

expilu edited this page Oct 3, 2026 · 1 revision

Auto mode

mode: 'auto' answers a question the way you would trust a good assistant: let the fast instinct answer, and only when it does not seem sure, hand the same question to deliberate reasoning. It is the mode for code that stays online — the fast path stays in the milliseconds, and the seconds of deliberation are spent exactly where they change an answer. The two passes are Under the hood: choice, score and noul all support it, whichever mode you pick.

How it works

  1. The question runs System 1 first: one forward pass, one generated token, milliseconds — the fast answer that is explained in Under the hood.
  2. The answer's confidence is read. For choice and score that is the returned confidence (flat distribution → low, single peak → high); for noul — which has no confidence field, because its single yes/no value already describes everything — the equivalent is max(noul, 1 − noul): a split verdict is a low-confidence verdict.
  3. If the confidence is below the threshold, the question escalates to System 2: the model deliberates and rates every candidate in the Under the hood, and that answer is returned as-is, whatever its own confidence.
  4. Otherwise the System 1 answer stands and System 2 is never called.

The escalation happens once. A still-doubtful System 2 answer is returned visibly low-confident — your code decides what to do with it (the same thresholding a human sees in Support ticket routing) — but the library never loops, so a cost ceiling is a design guarantee, not a hope. Since all three primitives return the same answer shapes whatever mode ran (choice's walkthrough shows the fields), your code does not need to know which pass won.

The confidence readings on both sides of the threshold are themselves observable under debug: true (Under the hood documents what every layer logs), as one gray line per question:

[smart-decisions:auto] confidence 0.40 < threshold 0.70 — escalating to System 2
[smart-decisions:auto] confidence 0.93 ≥ threshold 0.70 — keeping the System 1 answer

The threshold

autoModeThreshold is a question field; it is only valid in auto mode (set it anywhere else and the mode-partitioned question types — or a loud runtime error for JS callers — stop it). It defaults to 0.7.

Read it as "how unsure can a fast answer be before we pay for deliberation":

  • Higher (0.8–0.9): escalate more. More System 2 calls, more wall time and tokens; use it when a wrong verdict is expensive and deliberation works for your traffic.
  • Lower (0.4–0.6): escalate rarely. Only the signs that the System 1 distribution is flat get escalated; use it when latency dominates and being approximately right is acceptable.

The same number applies to every question of that call, whatever the primitive — the confidence semantics are shared (Under the hood documents how it is read on each primitive), so the one knob means the same for choice(), score() and noul().

What it costs

Almost nothing when System 1 is sure. Measured on the reference server (llama.cpp, Qwen3.5-4B on an aging RTX 2080): noul questions whose verdict leaned clearly never escalated in the latency benchmark — the mode clocked the same milliseconds as plain System 1 for the whole run. A score question whose fast answers sat below the 0.7 threshold escalated on essentially every round, and paid one deliberate pass each time — the seconds only where the model was genuinely split. That asymmetry is the point.

Deliberation's own cost shows up in System 2, not here: the Under the hood, and how a caller's timeoutMs protects against a stalled reasoning loop.

When to use it

  • Long-running classification flows with a human or fallback path downstream, like the Support ticket routing — the low-confidence System 2 answer replaces some of the human escalations, without paying deliberation on every question.
  • Traffic mixes where most questions are easy: confident verdicts never touch System 2, so the fast path stays in the milliseconds (choice is the typical traffic).
  • Whenever you would otherwise write "if confidence < X then run the careful prompt again by hand" — that is exactly this mode, with the retry-free, once-only cost rule built in.

For one-shot scripts where latency is air, plain System 1 or System 2 with the knob you want is simpler reasoning than auto.

Example

const answer = await choice({
  model,
  mode: 'auto',
  autoModeThreshold: 0.7, // the default, shown here to make it discoverable
  state: `Customer support ticket:\n"${ticket}"`,
  instructions: 'Classify the ticket into the right department',
  criteria,
});
// answer is a ChoiceAnswer either way — your code does not know which pass won,
// and 'debug: true' shows it in the escalation line.

The ticket-routing-auto.ts example in the repo runs three tickets through this exact shape; Examples links it and the other runnable ones.

See also

choice, score and noul for the primitives and the question fields, Under the hood for everything under both passes, Support ticket routing for the confidence story in a worked example.

Clone this wiki locally