Repository navigation
Under the hood
How the library turns a prompt for a tiny answer into a probability distribution over your options and a confidence, and what every one of the returned numbers is.
An LLM always computes a probability distribution over every possible token before it answers. That process is there whether you actually need a sentence down the line or not; normal chat completions requests ask for the sampled output and the API carries the distribution around anyway. The trick is to request the distribution itself:
- The model is asked a single question and reads the logprobs of the answer's first (and only) generated token. The generated token itself is irrelevant and gets discarded. What matters is the letter and digit logits making up the probability distribution over your options.
- The distribution over your options is exactly this kind of readout: the probability of the answer token being each of the letters or digits that stand for every option.
That single generated token is the entire decision: no deliberation, no chain of thought, no structured output parse (those are System 2's job.
Every System 1 question goes out as a plain non streaming POST to {apiBaseUrl}/chat/completions, carrying the prompt plus a set of settings tuned to make the single generated token the answer itself:
-
max_tokens: 1(the answer is a single token, one single pass of the model) -
temperature: 0(greedy: always the most likely symbol) -
logprobs: true, top_logprobs: 20(the report window of candidates for that one token) - thinking and reasoning disabled (
chat_template_kwargs: { enable_thinking: false }), so the one token isn't wasted on a think tag
The top_logprobs number is deliberately portable rather than generous: 20 is the window that OpenAI and OpenRouter cap at, vLLM's server default --max-logprobs is also 20, and llama.cpp accepts up to 50. At 20 the request is safe everywhere logprobs exist, and realistic option counts still fit: symbols that fall outside the window simply contribute 0 to their option's probability.
The prompt itself changes per question type only in the list and the closing ask. For choice(), each option is rendered as A: key — description, and the prompt closes with "Answer with exactly one letter (A, B, C). Reply with that single letter and nothing else."
The window that comes back is not options: it is the top tokens for one generation, which often include half spellings, spaces, punctuation. Each question type defines how a raw token folds into a candidate symbol:
- choice folds any case or spacing variant of a letter onto the bare uppercase letter: " B", "b" and "B." are all the same candidate.
- score takes an exact single digit and rejects everything else, so "10" or "1." or "three" cannot blur a neighboring level's probability.
- noul takes the first alphabetic run of the token, lowercased, and matches the words y, yes, yeah and yep on the yes side, and n, no, nope and nah on the no side. Anything else contributes nothing.
From there the readout is shared: for every symbol the best variant probability survives, all others contribute 0, and the symbol probabilities are normalized so the distribution sums to 1.
The winning option that comes back in answer.choice is whichever letter (folded, normalized) held the highest probability; what lands in your code is your own option name, never the letter.
confidence (present on choice() and score()) is the shannon entropy of the normalized distribution, divided by the entropy of the uniform one, complemented against 1:
H = -Σ p·ln(p) / ln(n), and confidence = 1 - H.
Read as a decision maker: an even spread across the options is maximum entropy, the model has no clear instinct, and the confidence nears or hits 0; a single peak sends it towards 1. A concrete trio worth having in your head, with two options: half and half is 0, all on one side is 1, 0.9 against 0.1 is about 0.56.
There is no confidence on noul(): with only two outcomes, the single yes/no probability describes the whole judgment, and a value near 0.5 already is the "no clear read" signal in itself.
When no candidate symbol at all shows up in the top logprobs window, the distribution falls back to uniform rather than empty: an all zero distribution would fake a mysterious "100 percent not this" answer, whereas a uniform one is useless in a visible, correct way. Concretely:
- On score: every level ties, the score lands on the exact spectrum midpoint and the confidence reads 0.
- On noul: the value lands on exactly 0.5, the visible neutral for a binary judgment.
- On choice: every option ties, the argmax falls on the first entry and the confidence reads 0, so a confidence threshold in your code catches the case.
System 2 has a twin case with the same fallback (a reply whose ratings are all zero reads uniform), which the ratings section below covers.
The companion rule on the transport side: a response that carries no logprobs at all is an error, not a fallback. There is no signal to read, and a made up uniform answer would silently lie.
System 2 asks the model the same question with room to deliberate and expecting a structured output (an object with certain shape) that the model must generate correctly. The prompt carries the state, your instructions, the candidate list, the schema itself (deliberately redundant with the response_format field, so schema-less inference engines still know the shape), and a closing instruction: rate each candidate with an integer from 0 (does not fit at all) to 10 (fits perfectly), and reply with a single JSON object.
The structured answer is minimal on purpose: candidates as keys, ratings as values, nothing else. There is no separately-asked "decision" field, the verdict is the argmax of the ratings, so the winner and the distribution can never disagree.
Where System 1 reads its probabilities from the token window, System 2 normalizes the ratings: each candidate's probability is its share of the total ratings mass. A deliberate 0 is information (unlike a symbol missing from a logprob window — the model weighed the candidate and ruled it unfit). The two-step mapping, per question type:
- choice — the ratings over option names, normalized, argmax wins (ties → the first option, as in System 1).
-
score — the ratings over level numbers
0..n, normalized; the position on the spectrum is the same weighted mean. -
noul — the ratio of the two ratings:
P(yes) = true / (true + false). Both rated 0 → uniform → exactly 0.5, the visible "no clear read" neutral.
confidence is the entropy of the normalized distribution divided by the uniform entropy, complemented against 1, the same reading as System 1: an even spread is 0, a single peak is 1, 0.9-against-0.1 is about 0.56.
The deliberate request asks for:
-
response_format: {type: 'json_schema', json_schema: {..., strict: true}}— the ratings schema, written strict-compatible (every property inrequired,additionalProperties: false) so the same document works on hosted OpenAI and on engines that convert schemas to grammars (llama.cpp → GBNF, vLLM → xgrammar). -
temperature: 0— deterministic verdicts. -
No
max_tokens— server defaults carry it, so a reasoning model's deliberation budget is not starved by our ceiling; reasoning tokens count against any budget you do set. - the reasoning toggle of your question's
thinkingfield (see below).
The reply is parsed, validated and, if anything is wrong, retried with the rejection fed back to the model, bounded by the question's maxRetries. This runs even when the engine enforced response_format, because it can't always:
- llama.cpp older builds ignored the grammar while thinking was enabled (#20345; fixed in current builds),
- engines without schema support echo whatever the template yields,
- reasoning parsers differ: thinking may arrive as a separate
reasoning_contentfield (llama.cpp, vLLM, DeepSeek, Ollama), areasoningfield (OpenRouter), or inline think blocks insidecontentitself, which the library strips before parsing, - reasoning tokens eat the token budget, so a truncated reply can arrive empty;
finish_reason: 'length'is treated as a rejection, not a crash.
An engine's structural shortcomings cost attempts, never correctness. After the budget is spent the error carries the last rejection reason, so the failure is diagnosable from the outside.
-
thinking: false— the library sendsreasoning_effort: 'none', the OpenAI-standard disable. Verified live on llama.cpp; documented for OpenAI, OpenRouter, Ollama's OpenAI-compatible endpoint and vLLM. -
thinking: true/ omitted — nothing is sent: every reasoning model thinks by default, and engines that need a different enable key take it throughextraBody.
The cost difference is real: thinking on takes much longer to complete a response, while hopefully is more accurate.
'auto' runs System 1 first — milliseconds, one token — and escalates once to System 2 when the answer's confidence is below autoModeThreshold (default 0.7). For noul, whose single value is its own confidence, the equivalent is max(noul, 1 − noul): a split verdict is a low-confidence verdict.
The escalated answer is returned as-is, whatever its own confidence: one deliberate pass, never a loop. Confident System 1 answers never touch System 2 — fast path stays milliseconds, and only genuinely close verdicts pay the seconds. The full article — threshold tuning, cost shape, when to use it — is Auto mode.
debug: true on a question (both modes) logs the internals to stderr — prompt, wire request, raw responses, every retry and rejection reason, the 'auto' escalation decision (confidence vs. threshold: which pass answered, and why), and in System 2 the model's reasoning text. Never logged: credentials and URLs' query parameters.
Documentation
Examples
Reference