Releases: solvi-ai/solvi
Release list
solvi 1.0.1
Fixed
- A compact record now keeps the answer a hard check's
then=function gave, and replay checks it. In 1.0.0 a
compact record (record="compact"or the compact part of"sample:N") left out the trace record of kindthen
(namethen:<question>) that athen=function of facts writes, soreplay_allreported every decision such a
function answered asnot_kept("not verified") instead of checking it — 51 of 100 compact decisions in the
playground's compact-journal demo. Compact records now keep it whole, like a model step (its value, error and the
hashes of the facts it read), andreplay_all/store.rederivere-run the function on the re-computed facts and
compare: an edited value, or a function that now gives another answer, is a mismatch atthen:<question>. A kept
record that does not match its step hash in the record is reported asintegrity(it wasnot_kept). Athenrecord
costs about 0.4 KB in a compact record (the demo: 1.8 → 2.0 KB a decision);benchmarks/journal_size.pydoes not
change (its workloads have nothen=functions). Full records do not change. Compact records written by 1.0.0
still load, verify and replay as before: they do not hold thethenrecord, so those decisions staynot_kept
(tests/fixtures/store_1_0_0).
solvi 1.0.0 — two levels, knowledge and agents
1.0 makes solvi two levels. The high level is ready systems you configure: solvi.build (decisions from
labelled examples with a promise on the errors), solvi.Agent (new: acting in an environment), solvi.Guard (an
agent's tool calls) and solvi.Knowledge (new: what a system knows and from whom — facts with their sources, exact
retraction, an agenda of goals and gates, an action model learned from outcomes). The low level, solvi.core, is
the building blocks they are made of, with every extension point exported, documented and covered by a conformance
check. What works but has not shown a measured gain is moved to solvi.experimental, marked in every decision
that uses it, with a deadline (1.2) to graduate or go. Field reports from building on solvi brought quotes matched on
a normalized view, quotes from several labelled sources, then= that wires its hard check and can compute the answer,
res.checks, a compact journal, budgets for generation and refinement, and outcomes as labels. Python 3.11 or newer.
Every 0.9 import path keeps working through 1.0.x with a warning, and solvi migrate rewrites your code. What was
tried for this release and did not meet its bar is listed at the end.
Breaking changes and migration
If your code imports only the 23 names solvi still exports (below), it runs unchanged. Every other 0.9 import path
still works in 1.0.x: it imports the very same module and warns (SolviDeprecationWarning) with the path to use;
1.1 removes the old paths. Run solvi migrate PATH once (below), then your tests with
-W error::solvi.SolviDeprecationWarning. What breaks now, without a warning period:
- Python 3.11 or newer (see Changed for the fingerprints this fixes).
- The removed modules (table under Removed) raise ModuleNotFoundError.
- The methods moved off the classes (table below) raise an AttributeError that names the new call.
then=wires its hard check. A hard check whosethennames a question now runs in that question's flow by
itself;requires=is no longer needed for it, and other questions' flows do not change. Before, forgetting
requireslet the question be answered as if the check had passed. A strategist of your own that leaves such a check
out is refused when theSystemis built;solvi check(then_not_in_flow) keeps reporting it for a strategist
swapped in afterwards. A system that relied on the check not running for that question now gets the forced answer
when it fails.
The new layout: solvi's modules moved into packages that say what you can rely on — solvi and solvi.solutions
are the ready-made systems, solvi.core and its areas (solvi.core.types, solvi.core.runtime, ...) are the
building blocks they are made of, and solvi.experimental holds what may still change.
solvi migrate PATHrewrites your code (.pyand.mdfiles) to the new paths: imports,from solvi import storage, dotted paths in strings such asmonkeypatch.setattr("solvi.llm.urlopen", ...)or
"solvi.hooks:rules_system".solvi migrate PATH --checkchanges nothing and exits 1 when a file would change.- Stored decisions, calibration files and fingerprints do not change: a fingerprint records the 0.9 module of a moved
class or function (the tablesolvi._deprecate.MOVED), so decisions stored by 0.7–0.9 replay, and a store written by
1.0 is read by 0.9 tools the same way. A storedmodule:qualnamethat names a 0.9 module loads without a warning. - What
solviexports: 23 names (__all__) — the entry pointsbuild(solvi.solutions.decisions.build, was
solvi.auto.build),Agent(new,solvi.solutions.agent),Guard(solvi.solutions.guard.Guard, was
solvi.agents.Guard) andKnowledge(new,solvi.solutions.knowledge),Budget, and the shared vocabulary
Catalog,Question,Answer,System,Response,Quote,Claim,Decision,Fail,Unknown,Span,
Maybe,Rank,Estimate,Scale,Bins,SolviDeprecationWarning,ExperimentalWarning. The other 0.9 names (JSONLStorageand the other stores,
TraceStorage,Trace,Record,Result,MISSING,Shadow,AnswerType,NotStated,FactTypeError) are
imported from their modules (table below);from solvi import JSONLStorageworks in 1.0.x with a warning. solvi.auto.AutoSystemis nowsolvi.solutions.decisions.DecisionSystem(whatsolvi.buildreturns; the old name
works in 1.0.x with a warning).build(slow=...)no longer compiles a written specification itself (that would make
it import the experimental compiler): compile it first (solvi.experimental.compile.compile_spec) and pass the
result;writer=andinputs=raise a TypeError saying so.
Where each module went
| you imported (0.9) | import now (1.0) | level |
|---|---|---|
solvi.agents |
solvi.solutions.guard |
high level: ready to use |
solvi.agents.confirm |
solvi.solutions.guard.confirm |
high level: ready to use |
solvi.agents.guard |
solvi.solutions.guard |
high level: ready to use |
solvi.agents.intents |
solvi.solutions.guard.intents |
high level: ready to use |
solvi.agents.mcp |
solvi.experimental.mcp |
experimental: may change; removed in 1.2 unless it graduates |
solvi.agree |
solvi.core.slow.agree |
low level: building blocks |
solvi.audit |
solvi.core.store.audit |
low level: building blocks |
solvi.auto |
solvi.solutions.decisions |
high level: ready to use |
solvi.calibfile |
solvi.core.calibfile |
low level: building blocks |
solvi.calibration |
solvi.core.calibration |
low level: building blocks |
solvi.charts |
solvi.experimental.charts |
experimental: may change; removed in 1.2 unless it graduates |
solvi.charts.check |
solvi.experimental.charts.check |
experimental: may change; removed in 1.2 unless it graduates |
solvi.charts.propose |
solvi.experimental.charts.propose |
experimental: may change; removed in 1.2 unless it graduates |
solvi.charts.render |
solvi.experimental.charts.render |
experimental: may change; removed in 1.2 unless it graduates |
solvi.charts.spec |
solvi.experimental.charts.spec |
experimental: may change; removed in 1.2 unless it graduates |
solvi.command |
solvi._command |
internal |
solvi.compile |
solvi.experimental.compile |
experimental: may change; removed in 1.2 unless it graduates |
solvi.costs |
solvi.core.costs |
low level: building blocks |
solvi.counterfactual |
solvi.experimental.counterfactual |
experimental: may change; removed in 1.2 unless it graduates |
solvi.decide |
solvi.core.deciders |
low level: building blocks |
solvi.decide.adapt |
solvi.core.deciders.adapt |
low level: building blocks |
solvi.decide.backends |
solvi.core.deciders.backends |
low level: building blocks |
solvi.decide.capabilities |
solvi.core.deciders.capabilities |
low level: building blocks |
solvi.decide.gate |
solvi.core.deciders.gate |
low level: building blocks |
solvi.decide.kinds |
solvi.core.deciders.kinds |
low level: building blocks |
solvi.decide.model |
solvi.core.deciders.model |
low level: building blocks |
solvi.decide.part |
solvi.core.deciders.part |
low level: building blocks |
solvi.decide.state |
solvi.core.deciders.state |
low level: building blocks |
solvi.decide.wire |
solvi.core.deciders.wire |
low level: building blocks |
solvi.diff |
solvi.core.store.diff |
low level: building blocks |
solvi.dispatch |
solvi.core.dispatch |
low level: building blocks |
solvi.drift |
solvi.core.guarantees.drift |
low level: building blocks |
solvi.episode |
solvi.core.knowledge.episodes |
low level: building blocks |
solvi.extract_long |
solvi.core.extract |
low level: building blocks |
solvi.generate |
solvi.core.slow.generate |
low level: building blocks |
solvi.guarantee |
solvi.core.guarantees.guarantee |
low level: building blocks |
solvi.heads |
solvi.core.deciders.heads |
low level: building blocks |
solvi.honesty |
solvi.testing.honesty |
high level: ready to use |
solvi.hooks |
solvi.experimental.hooks |
experimental: may change; removed in 1.2 unless it graduates |
solvi.i18n |
solvi.core._i18n |
internal |
solvi.inputs |
solvi.core._inputs |
internal |
solvi.learning |
solvi.experimental.learning |
experimental: may change; removed in 1.2 unless it graduates |
solvi.llm |
solvi.core.deciders.llm |
low level: building blocks |
solvi.loader |
solvi._loader |
internal |
solvi.longdoc |
solvi.core.deciders.longdoc |
low level: building blocks |
solvi.lora |
solvi.experimental.lora |
experimental: may change; removed in 1.2 unless it graduates |
solvi.memory |
solvi.core.knowledge.memory |
low level: building blocks |
solvi.multi |
solvi.core.deciders.combine |
low level: building blocks |
solvi.openset |
solvi.core.guarantees.openset |
low level: building blocks |
solvi.perturb |
solvi.core.deciders.perturb |
low level: building blocks |
solvi.primitives |
solvi.core.primitives |
low level: building blocks |
solvi.provenance |
solvi.core.provenance |
low level: building blocks |
solvi.refine |
solvi.core.slow.refine |
low level: building blocks |
solvi.remote |
solvi.core.deciders._remote |
internal |
solvi.report |
solvi.core.store.report |
low level: building blocks |
solvi.rulelist |
solvi.core.deciders.rulelist |
low level: building blocks |
solvi.runtime |
solvi.core.runtime |
low level: building blocks |
solvi.sandbox |
solvi.experimental.compile.sandbox |
experimental: may change; removed in 1.2 unless it graduates |
solvi.scaffold |
solvi.cli._scaffold |
internal |
solvi.schema |
solvi.core.schema |
low level: building blocks |
solvi.search |
`solvi.cor... |
solvi 0.9.0 — System 1 and System 2
0.9 is about two ways of deciding in one system, in Kahneman's sense: System 1, fast and cheap — rules, checks, a
fitted head, a model under a guarantee — answers when it is sure; System 2, slow and deliberate — an LLM, a re-ask
loop, a search — is woken when System 1 is unsure or surprised, and a person gets what neither can answer. A
dispatcher puts both in one recorded, replayable decision within a budget and can be calibrated on the inputs System 1
hands over; a system report tells the owner who answered, at what cost, and whether the promise held; System 2 can also
write what System 1 then runs fast — a policy text compiled into catalog parts, with a person settling what the drafts
dispute, and into an agent guard. Search over candidates runs about twice as fast, the nine-task benchmark stand reruns
in CI, and a showcase on the world map of Pokémon Red puts the pieces together. The names 0.8 renamed and kept with a
warning are removed. What was tried for this release and did not meet its bar is listed at the end.
Breaking changes
The old names that 0.8 kept working with a SolviDeprecationWarning are gone. An old keyword now raises a TypeError
and an old attribute or method an AttributeError; both say "X was renamed in 0.8 and removed in 0.9: use Y". The five
old modules are gone (importing one raises ModuleNotFoundError). If your code ran under 0.8 without a
SolviDeprecationWarning (pytest -W error::solvi.SolviDeprecationWarning finds them all), it runs under 0.9
unchanged. solvi.SolviDeprecationWarning itself stays, for later renames.
Removed (old → what to use):
| removed | use |
|---|---|
module solvi.fast |
solvi.heads |
module solvi.learned |
solvi.costs (CostBook, MeasuredCosts) and solvi.strategist (OrderModel, ProducerPolicy, Binary, scalar_row) |
module solvi.rules |
solvi.rulelist |
module solvi.strategy_model |
solvi.segment_model |
module solvi.extract_model (SpanExtractor) |
solvi.extract_long.LongSpanExtractor |
System(inputs=), system.inputs |
System(input_model=), system.input_model |
System(costs=), system.costs |
System(cost_policy=), system.cost_book |
System(journal=path) |
System(storage=JSONLStorage(path)) or storage="file.jsonl" |
ask(names=), aask(names=), Service.ask(names=) / aask(names=), Shadow.ask(names=) |
questions= |
Question(checkpoints=), q.checkpoints, part.question(cat, checkpoints=) (also on a combination) |
requires=, q.requires |
res.computed_state, res.computed_state_text(lang) |
res.state_text(lang) |
res.textin |
res.read |
system.teach(source=), store.save_correction(source=) |
label_source= |
system.learn_rule(facts=) |
features= |
system.calibrate(question, states, truth) |
system.calibrate(question, [(state, answer), ...]) |
system.safeguard_report() |
system.safeguard_summary() |
shadow.report() |
shadow.summary() |
system.fit_fast(...) |
system.fit(..., select=False) |
System.learning(harvest_rules=) (ignored in 0.8) |
— (teach rule outcomes with label_source="rule") |
trace.value(name) |
res.values[name]; a given fact: trace.init[name] |
answer_type.rank(v) |
answer_type.options.index(v) |
model.decision(escalate_below=, act_threshold=, target_error=, unknown=) and the same in decide, adapt, fit, teach, reset, decisions, questions |
min_confidence=, min_act=, max_error=, not_stated= |
a pydantic field's json_schema_extra keys "escalate_below", "act_threshold", "target_error" (now a ValueError) |
"min_confidence", "min_act", "max_error" |
model.has_unknown |
model.has_not_stated |
model.long_len, part.long_len |
max_len_long |
act_guard(risk=) (a part and a combination), adapt_lora(risk=), CorrectionMemory.calibrate(risk=), Guard.calibrate_authorizer(risk=) |
max_risk= |
calibrate_for(error=) (a part and a combination) |
max_error= |
the result keys "coverage" and "target_error" of calibrate_for |
"answered", "max_error" |
the result key "calls" of a combination's act_guard; combination.usage() and its key "per_question" |
"calls_per_question"; combination.calls() |
a combination's decide(x=), teach(x=), adapt(inputs=) |
text=, text=, texts= |
FastHead.update(row, answer), Binary.observe(row, y) |
teach(...) |
FastHead.cv_acc, Head.loo_acc (each head's number under the other's name) |
FastHead.loo_acc, Head.cv_acc |
ModelStrategist() without a model |
CostStrategist() (ModelStrategist now needs its model) |
CostStrategist(fallback=, fallbacks=) / ModelStrategist(...), .fallback, .fallbacks |
on_failure=, keep_alternatives= |
MultiSpanExtractor.fit(docs, spans), predict_doc(text) |
fit([(text, spans), ...]), predict(text[, field]) |
store.forget(fact, value) |
store.where_is(fact, value) |
JSONLStorage(path, catalog=) (every backend), store.get(id, catalog=) |
system= |
store.query(catalog=fingerprint) |
query(catalog_fp=) |
solvi.testing.check(system, case, state) |
solvi.testing.run_case(...) |
a honesty case's "gold" (now a ValueError); a solvi test case's "gold" (now reported as a problem of the case) |
"expected" |
solvi hook ... --model M (exits 1 with what to do, so an old installed hook blocks nothing), $SOLVI_HOOK_MODEL |
--decider M (reinstall the hooks), $SOLVI_HOOK_DECIDER |
solvi.serve.Guard (the ASGI middleware) |
solvi.serve.AccessGuard |
agents.Guard(facts=[names]) |
Guard(fact_names=[names]) |
the agent adapters' declare= (guard_tool, guard_wrappers, guarded_tool_node, GuardedToolset) |
auto_declare= |
run_proxy(context_messages=, context_chars=), solvi.agents.mcp.Proxy(...) the same |
max_messages=, max_chars= |
solvi.llm._error_text (private) |
solvi.llm.error_text |
Stores, calibration files and fingerprints are unchanged by these removals: the stored keys stay as they were (a
question still hashes its requires under the key checkpoints, a part still stores escalate_below). A test loads,
verifies and replays stores written by 0.8.0 and by 0.7.1 (tests/test_store_0_8_0.py, tests/test_store_0_7_1.py).
One fingerprint changes once: the code fingerprint of a catalog (its rules, functions and checks) no longer depends on
the Python version (3.12 and 3.13 print a function's syntax tree differently), so it differs from the one 0.8 computed
for the same code. Decisions stored by 0.8 still load, verify and replay; anything that compares a stored catalog
fingerprint with the current one (store.query(catalog_fp=...), a version check of your own) sees the catalog as
changed once. Model and question fingerprints, and a decision part's, are unchanged (tests/test_store_0_8_0.py).
New
- One entry point:
solvi.auto.build(preview). The question, labelled examples, the promise (max_risk=or
max_error=) and, optionally, a slow path (an LLM, a decision part, a System, a function or a compiled
specification), a budget and a store in; a ready System 1 and dispatcher out, withask(),report()and
explain()— a plain account of what System 1 is, the signal its guarantee reads, how the examples were split, the
promise and its threshold, who answers each slice it hands over and what is not covered. It composes what the library
has: the catalog's rule, your own fitted part (learner=) or a head fitted on the computed facts; the act
probability, the confidence or the computed number that best separates right from wrong; an open-set gate for a
choice among more than two options;Dispatcher.calibrateon examples System 1's guarantee did not see — only when
System 1 actually hands the slow path a slice, else those examples calibrate System 1. Measured on four tasks of the
stand against the hand-written setups: as good or better with every promise kept on eval, at a quarter to half of
the code. Limits: it lands closer to the promised level than the hand-written setups (risk 9.3% of 10%, 0.9% of 1%;
error 4.9% of 5% after a shift), it rarely gives the slow path anything (inputs unlike the examples go to a person,
and on the hard slice System 1's own guess was often the better answer), and its drift flag came later than the
hand-written one (191 vs 68 requests after a shift). Example:examples/24_one_entry_point.py. - Who answers:
solvi.dispatch(experimental).Dispatcher(system1, SlowPath(system2), ...)asks System 1
first; when its own signals say its answer cannot be given alone (below its guarantee, outside the open-set gate,
an abstention, a broken constraint, low agreement), the slow path answers or checks — a System, a re-ask loop
(propose=,solvi.refine) or a search (space=,solvi.search) — and when that cannot answer either, or the
budget (Budget(usd=, calls=, ms=), per decision and in total) is spent, a person gets the input with both
candidates and the reasons, never a guess. Drift and a sampled share (supervise=) make the slow path check System
1's answers instead. Every decision is one hash-chained record with its cost (dollars from the recorded tokens),
andd.replay(res)/d.replay_all()re-check it without calling a model. Guide: "Who answers". - Calibrating who answers on the hard slice:
Dispatcher.calibrate(examples, max_risk=...)measures, on each
slice of the inputs System 1 hands over, System 1's own would-be answer, the slow path's, and the slow path's when it
agrees with System 1, and picks one answerer per slice with a threshold — or a person — so that all answers given
alone keep the promise. The inputs handed over are the hard ones, where an LLM that is rarely wrong on average can be
wrong often and System 1's own guess can be the better answer; calibrate measures that instead of assum...
solvi 0.8.0 — one name per concept, any model first
0.8 gives every concept one name, makes the surface smaller and the decider protocol one, puts any model first (an LLM
through solvi.llm, a decision service, or a local checkpoint for offline use), and fixes what an independent audit of
0.7 found. Most renamed names still work in 0.8 with a warning that names the new one; they go in 0.9. The warning is
solvi.SolviDeprecationWarning, a FutureWarning, so Python shows it to you (once per old name per process) in your
own code too, not only in tests and __main__ as it would a DeprecationWarning.
Stores, calibration files and fingerprints written by 0.7.1 load, verify and replay unchanged (a test replays stores
written by 0.7.1).
Breaking changes
Removed or changed without a working alias:
- Every
Systemoption afterquestionsis keyword-only:System(cat, qs, "file.jsonl")is a TypeError —
System(cat, qs, storage="file.jsonl"). Everyask/aaskoption after the questions is keyword-only too:
ask(state, ["q"], 4)→ask(state, ["q"], workers=4). act_guard(of a part and of a Cascade / Vote / Route),calibrate_for,adapt_lora,Guard.calibrate_authorizer:
every option after the examples is keyword-only —act_guard(examples, 0.1)→act_guard(examples, max_risk=0.1).
systemone(url, model, key, 30.0)→systemone(url, model, key, timeout=30.0).- The octonion signature left the package:
sign(obj, "octonion"), reading an octonion signature andsolvi verify --sign --alg octonionraise a ValueError pointing tobenchmarks/octonion_signature.py(the default syndrome code
locates the same changes, 4x smaller, ~60x faster).solvi verifyhas no--algoption. model.decision(...)raises for an option the question's kind does not use, where it accepted and ignored it (some
changed the part's fingerprint):k=off a ranking,score_value=off a score question,bins=/unit=/
coverage=off a number question,other=on a score or yes/no question,min_margin=on a multi-label one,top_k=
/rerank=withoutlong=,min_act=/max_error=on a checkpoint without an act head,kind=that contradicts
multi=True.decisions(schema, fields=[...])raises for a field the schema does not have.- A wrong key, model or URL for a System One service (HTTP 401, 403, 404, another 4xx except 400 / 413 / 422) raises
SystemOneError, assolvi.llmraisesLLMError(both aresolvi.remote.RemoteError); it used to escalate every
decision, andsolvi models checkthen reported "answered alone 0.0%" and exited 0. scorer.usageof an LLM decider andextra["llm"]["usage"]countinput_tokens/output_tokens/
reasoning_tokens, as a System One decider does (they wereprompt_tokens/completion_tokens).guarantee["signal"]of a Cascade / Vote / Route is a name —"shared","shared-rank"— as a part's is ("act",
"confidence"); it was a sentence. Fingerprints and calibration files of 0.7 stay valid.- Escalation texts of remote models: "the LLM server refused the request: HTTP 400 — ..." (was "invalid input for the
endpoint: ..."), "the LLM server did not answer after 3 attempts: ..." (was "no answer from ... after 3 attempts"). - Command lines that pass an option their mode does not read are refused (status 2) instead of being ignored:
solvi servein HTTP,--mcpand--guardmodes;solvi ask --reportwith--json,--auditor--lang;--backend/
--api-keywithout--decider;solvi report --idwith period filters.POST /askandPOST /ask_textrefuse a
body key they do not know (422) — a misspelled"question"used to ask every question. solvi ask --text --jsonprints the text read as"read"(was"textin"), asPOST /ask_textdoes.- Removed, nothing called them:
Catalog.producer,strategy.fact_of,Selection.expanded,serve.RequestTimeout,
SystemOneScorer.question,ModelStrategist(max_expand=)/search(max_expand=). - New in this release and renamed before it is published (no alias):
solvi.many.fit→fits(its record key too),
DriftMonitor.calibrate→set_reference,EpisodeView.revisits(kind, key)→revisits(key, kind)with the key
required,Episode.note(kind, key, value)→note(kind, key),episode.Chooser(escalate_below=)→
min_confidence=;System.guarantee,solvi.guarantee.calibrateandOpenSetGate.calibratetakemax_risk=/
max_error=only.
Renamed, the old name works in 0.8 with a warning (old → new). Every old name warns once per process with
solvi.SolviDeprecationWarning (a FutureWarning, shown by Python's default filters): "X is deprecated since 0.8 and
will be removed in 0.9: use Y". To find them all at once, run your tests with -W error::solvi.SolviDeprecationWarning;
to silence them, warnings.filterwarnings("ignore", category=solvi.SolviDeprecationWarning).
| 0.7 | 0.8 |
|---|---|
System(inputs=Model), system.inputs |
System(input_model=Model), system.input_model |
System(costs="measured"), system.costs |
System(cost_policy="measured"), system.cost_book |
System(journal=path) |
System(storage=JSONLStorage(path)) (or storage="file.jsonl") |
system.ask(state, names=...), aask(names=), Service.ask(names=), Shadow.ask(names=) |
questions= |
Question(checkpoints=[...]), q.checkpoints, part.question(cat, checkpoints=) |
requires=, q.requires |
res.computed_state, res.computed_state_text(lang) |
res.state_text(lang) |
res.textin |
res.read |
system.teach(..., source=), store.save_correction(..., source=) |
label_source= |
system.learn_rule(question, examples, facts=) |
features= |
system.calibrate(question, states, truth) |
system.calibrate(question, [(state, answer), ...]) |
system.safeguard_report(), shadow.report() |
safeguard_summary(), summary() |
Trace.value(name) |
res.values[name] (a given fact: trace.init[name]) |
AnswerType.rank(v) |
answer_type.options.index(v) |
model.decision(escalate_below=, act_threshold=, target_error=, unknown=) (and decide, decisions, questions, a field's json_schema_extra) |
min_confidence=, min_act=, max_error=, not_stated= |
model.has_unknown, model.long_len, part.long_len |
has_not_stated, max_len_long |
act_guard(risk=), adapt_lora(risk=), CorrectionMemory.calibrate(risk=), Guard.calibrate_authorizer(risk=) |
max_risk= |
calibrate_for(error=); its result's "coverage", "target_error" |
max_error=; "answered", "max_error" |
a combination's act_guard()["calls"], combination.usage() ("per_question") |
"calls_per_question", combination.calls() |
FastHead.update(row, answer), Binary.observe(row, y) |
teach(...) |
FastHead.cv_acc, Head.loo_acc (each head's number under the other's name) |
FastHead.loo_acc, Head.cv_acc |
ModelStrategist() without a model; fallback=, fallbacks= |
CostStrategist(); on_failure=, keep_alternatives= |
solvi.fast, solvi.learned, solvi.rules, solvi.strategy_model |
solvi.heads; solvi.costs + solvi.strategist; solvi.rulelist; solvi.segment_model |
solvi.extract_model.SpanExtractor |
solvi.extract_long.LongSpanExtractor |
MultiSpanExtractor.fit(docs, spans), predict_doc(text) |
fit([(text, spans), ...]), predict(text[, field]) |
store.forget(fact, value) (it never deleted anything) |
store.where_is(fact, value) |
JSONLStorage(path, catalog=) (all backends), store.get(id, catalog), store.query(catalog=fingerprint) |
system=; query(catalog_fp=) |
solvi.testing.check(system, case, state) |
run_case(...) |
a honesty case's "gold" |
"expected", as in solvi test cases |
solvi hook ... --model M, $SOLVI_HOOK_MODEL (hooks installed by 0.7 keep working) |
--decider M, $SOLVI_HOOK_DECIDER |
solvi.serve.Guard (the ASGI middleware) |
solvi.serve.AccessGuard |
agents.Guard(facts=[names]) |
Guard(fact_names=[names]) |
the adapters' declare=True |
auto_declare=True |
run_proxy(context_messages=, context_chars=) |
max_messages=, max_chars= |
System.learning(harvest_rules=True) |
— (it harvested nothing in the usual wiring; ignored) |
Structure: solvi.decide is a package of layers (kinds, state, wire, capabilities, backends, adapt,
gate, part, model) and re-exports every name it had; the combinations' base is public as
solvi.multi.Combination; ExperimentalWarning and the "not stated" option name live in solvi.core; the helpers
every command shares are in solvi.command, so no library module imports from solvi.cli; every module declares its
public names in __all__, and the API reference shows only those (a helper that is not in __all__ may change without
notice).
New
- Any model first. The README and the guide present the decider as whichever model you have: an LLM through
solvi.llm(the core install is enough), a System One service, or a local checkpoint such as solvi-base for offline
or cheap use — with solvi-base's model-card numbers where it is offered. solvi.remote: the client every remote model shares (solvi.llm,solvi.systemone,solvi.generate, the chart
proposer) — the endpoint, the key in the header only, retries with backoff, token counts under one set of names, and
one policy for HTTP errors: a wrong key, model or URL raises, a refused input escalates that decision, no answer is
retried and then escalates.- One decider protocol. A decision part and a Cascade / Vote / Route have the same public methods, with the same
signatures and result keys:act_guard(examples, *, max_risk, signal, groups, min_group, delta)(a combination's
result withcalls_per_question,cost,scaleandanswered_bybesides), and now on a combination too
calibrate_for(examples, *, max_error, signal, method, delta)(one shared threshold for...
solvi 0.7.1 — solvi behind a coding agent's hooks (preview), gallery for coding agents, System One hardened, benchmark vs LLMs
solvi behind a coding agent's hooks (preview)
solvi hook pre-edit --rules rules.toml: a PreToolUse hook for Claude Code's Edit, Write and MultiEdit. It reads the
proposed change from the hook's JSON, works out the added lines with their line numbers in the file after the edit,
and asks a small solvi System one question,edit∈ {allow, deny, ask}, whose hard checks are the rules whose path
globs match:forbid(regular expressions over the added lines),require(over the file after the edit),
forbid_callsandrequire_def(Python, from the parsed code), a rule with no checks (any change to these paths), and
a fuzzyquestiona decider answers when an added line matcheswhen. It answers "deny" with the rule, the lines and
the rule's reason (the agent reads it and can fix the change), "ask" (the user confirms), or nothing (Claude Code's own
permissions apply;--approveanswers an explicit allow). A fuzzy rule blocks only with a calibration file
(act_guard: P(answered alone and wrong) ≤ risk on labelled changes); without one its "yes" asks, and without a model
a triggered question asks. A check that cannot run, instruction-like text addressed to a reviewer in the added lines,
and a hook that fails all ask — never a silent allow.solvi hook pick-skill --skills-dir .claude/skills: a UserPromptSubmit hook that picks one skill from the skills'
names and descriptions (the words a prompt shares with each, weighted by rarity; or a decider's choice with--model)
and adds one line naming it asadditionalContext; silent on "none", a near tie or a slash command.--model: a local checkpoint (never downloaded), a System One service (systemone:URL#model;solvi serve --decider ... --model-name ...keeps a local model loaded), an OpenAI-compatible endpoint (llm:URL#model, key from the
environment) or your own decider. Calibrate a rule's question withsolvi calibrate solvi.hooks:rules_system RULE_answer labels.jsonl --risk 0.1($SOLVI_HOOK_RULES,$SOLVI_HOOK_MODEL) and name the file in the rule.- Every decision is stored with its trace (
.solvi/traces/hooks.jsonlby default, any TraceStorage with--store),
sosolvi verifyandsolvi reportwork on it; parallel hooks share one chain (a file lock), and the store opens from
its head, not by reading every record.solvi hook audit [ID]prints a stored decision's audit and replays it against
the current rules. solvi hook installmerges the hook entries into the project's.claude/settings.json(other hooks and settings
stay; its own are replaced, not doubled), writes the sample rules (solvi hook sample-rules: no secrets in source, no
employee data taken from the browser in app/api, reversible migrations, no eval or shell strings, a person for CI
workflows) and prints what it changed;solvi hook uninstallremoves exactly its entries.- Codex (preview):
--agent codexwrites.codex/hooks.json, reads theapply_patchenvelope, and answers in Codex's
dialect (no "ask": a deny that says a person must confirm). - Speed: without a model a hook call is one short process — 105–121 ms on a laptop (median), with a store of 3000
decisions. - examples/22_coding_agent_hooks.py: a session in a temporary project — a clean
edit, a rule broken, a comment that tries to talk past the rules, two prompts, the verified store and one audit.
Benchmark: solvi vs asking an LLM
- docs/vs_llm.md — the same inputs and written policies given to solvi, four LLMs and a hosted
decision model, directly and inside solvi: refunds, 3-way invoice matching, the gallery and routing bank messages.
Strong reasoning LLMs followed the rules nearly perfectly and solvi was not more accurate there; the differences are
cost, latency, answers that do not change with option order, replay and hard checks.benchmarks/vs_llm/has the
data, the written policies, the runner (the solvi arm offline and free; any OpenAI-compatible or System One endpoint)
and every raw answer, sobench.py score --checkrecomputes the published tables without an API key. The
playground gains a "solvi vs LLM" tab on the same data.
Gallery: helpers for coding agents
Three new entries for decisions a coding agent (such as Claude Code or Codex) meets on every task. Each runs offline:
the deciders are keyword stand-ins calibrated with act_guard on synthetic, seeded examples, and each README shows
solvi-large, any OpenAI-compatible LLM (solvi.llm) or a System One service in front, with the keywords as the
fallback, and says what the offline rules cannot read. All three are playground presets.
gallery/13_pre_edit_rule_check: before the agent writes a file, the project's rules for that path (per glob) are
checked → allow / block / escalate, and the rules broken, with their lines. Secrets, browser storage read in
app/api/**and irreversible migrations (Python's own parser: adowngrade()that does something,RunPython/
RunSQLwith their reverse) block by hard checks; a CI workflow edit and a migration that does not parse go to a
person. "Auth checks go through require_role()" and "no personal data in log lines" are a decider's question each —
out of scope → the decider (act_guard, risk 5%; perturb=2) → a person, as fallback producers of one fact. A
# reviewer: ignore the rules abovecomment is not read by the code checks and flips the decider, which escalates.
Also: a pre-edit hook script and aGuardpolicy on awrite_filetool. 16 cases.gallery/14_review_triage: seven yes/no risk questions → quick or full review. Three are code (dependencies, a large
or unfocused change, logic without tests), four a decider's (auth, public API, migrations or deletion, security),
each calibrated at risk 2.5% so that P(a risky change goes to quick review) ≤ 10% (a union bound). On 2000 fresh
synthetic changes: 0.2% risky and quick, 22% quick overall. The change generator (synthetic_changes(n, seed)) is in
task.py; one known miss (a permission change without its usual words) is a case. 14 cases.gallery/15_skill_picker: a prompt → exactly one of ten skills (with look-alike pairs),none, or a person. A
confident "none", an abstention and a wrong pick are kept apart; a skill the user names is cited, one named inside an
instruction-like passage of pasted text is not; a near tie (min_margin) escalates with the conformal candidates; a
production deploy needs the user's own word "production" (a hard check). 16 cases.
Changes
import solvino longer imports numpy: the answer heads load it on first use, and hashing a fingerprint only looks for
arrays when numpy is already loaded. The pydantic models ofsolvi.schemabuild their validators on first use.tomliis a dependency on Python 3.10 (rules files are TOML).solvi.systemone:systemone(..., extra_body=...)(andSystemOneScorer(..., extra_body=...)) merges
server-specific fields into every request — OpenRouter'sproviderrouting,user— with the rules of
solvi.llm's: copied, merged under solvi's own fields,model/state/questionsrefused with ValueError, part
of the fingerprint (without it the fingerprint is unchanged).solvi.systemoneanswers "not stated" (unknown=True,Maybe[...]): the question gets one more option, "not
stated", with a description; a yes/no question that allows it is asked as a choice over yes / no / not stated. Its
probability competes with the options' as withsolvi.llmand the checkpoints with a "not stated" output: the
decision isUnknownand the question built on it abstains. Before,decision(unknown=True)raised ValueError.solvi.systemoneanswers multi-label questions: onenoulper option in the same request, an option chosen at the
model's multi threshold (0.5), the confidence the least sure option's max(p, 1 − p) — what act_guard calibrates on;
with "not stated" allowed, one morenoulfor it.SystemOneScorer.questions(item)gives an item's questions;
question(item)still gives the one question of a single-question item. Spans and evidence stay refused.solvi.systemonerecords per decision, inextra["systemone"], the endpoint, the model name (served_bywhen the
service names another), the request'smsand, when the service reports them, itsusagetokens andcost— for
the whole request (questions: how many questions it answered); the scorer sums them inusageandcost.
costs="measured"already plans on each part's measured run time, the request included.- Behaviour change:
solvi.systemonehandles a service that fails assolvi.llmdoes, instead of raising:
network errors, timeouts, a broken connection (IncompleteRead, a reset), 408 / 409 / 429 and 5xx are retried
(retries=2,backoff=1.0s doubling), then the decision escalates ("the System One service did not answer after 3
attempts: ...") and is not cached, so the next ask tries again; another 4xx escalates at once with the service's error
text, a gateway's wrapped cause included (OpenRouter'serror.metadata.raw); a reply that breaks the contract (a
missing answer or probability) escalates ("invalid System One output — ..."). The API key is never in the reason. - Behaviour change: a System One model is
deterministic=Falseby default, assolvi.llm's: replay checks the
recorded output instead of calling the service again.systemone(..., deterministic=True)keeps the old re-run for
a local server whose output is reproducible.
Fixes
- The audit's guarantee line said "none for some decisions: their thresholds were not calibrated" when two decisions
behind one answer carried the same promise (identical promises were counted once against the number of decisions). - A decision escalated by
min_margin(a near tie) is now a low confidence ...
solvi 0.7.0 — text in, agent guard (preview), several models with an LLM stage, learning from corrections, long documents, LoRA (experimental)
The agent guard (solvi.agents) ships as a preview: its hard line is provenance (a value found only in a tool's
output never grounds an argument that must come from the user) and your policies; detecting injected instructions in
text is a heuristic second line. System.learning is experimental and off unless you call it; part.adapt_lora is
experimental too. Three code reviews and three adversarial passes ran before this release; their fixes are listed under
"Fixes before release".
Text in: entry points
system.entry_points(names=None): the questions as entry points — name, text and the typed input fields each one reads
(type, description, required), from the same schemas assolvi serve;ep.tool()is the function-calling form.solvi.textin.TextIn(system, decider, extractor=None, ...):read(text)→ aTextRead— the entry point the decider
picks (a choice over the entry points and their descriptions; escalates belowmin_confidence=0.6, on a near tie
min_margin=0.1or on the decider's act signal), and each input field read by span extraction with a quote and a
deterministic parser per type: numbers ("1,500.50", "1.5 million", "2k", "полтора миллиона"), dates ("2026-09-12",
"12.09.2026", "12 September", "12 сентября"; year-less and relative dates only withtoday=), enums by label or
synonym, booleans, strings (withpatterns=). A field isread,not_stated,unparsed,unsureorunsupported;
required fields not read are inread.missing, andread.clarify()asks for them — nothing is guessed.- Extractors: the decider's span pointer (
DeciderExtractor, when the checkpoint has one) orCueExtractor(deterministic
candidates of the field's type nearest after a cue word); any object withfind(text, FieldSpec) → [Quote]. system.ask_text(text | TextRead, decider=None, *, textin=None, question=None)(andaask_text): TextIn + ask in one
trace. The text is a given fact (request_text); the entry point (textin, provenancedecided) and each field
(textin:<field>, provenancequoted, with the extractor's fingerprint, the parser and its arguments) are hash-chained
records. The audit shows the fields as quoted by a model — never given, not in the deterministic share — and an answer's
confidence is at most the reading's. Replay re-checks each quote, re-parses it and checks the flow read that value. An
escalated entry point runs nothing: the likely questions abstain with guardescalated.res.textinis the TextRead.- A dialogue:
tin.update(read, next_message)reads the next turn over the whole dialogue and listschanges(old value,
new value, quote); "not A-10457 but A-10475" changes the field to the new value.
Guarding an agent's tool calls
solvi.agents.Guard: an agent proposes a tool call ({"name", "arguments"}— data, never code; OpenAI, LangChain,
Anthropic and MCP shapes are read byToolCall.parse) and solvi checks it as a proposal: the tool is in the catalog
(@guard.toolon typed functions,guard.declare(name, schema=...)for a pydantic model or a JSON schema), the
arguments validate against its types (unknown arguments are errors), theground=arguments are quoted from the
conversation (strings literally, numbers as number tokens, lists item by item;ground_from=the roles allowed — never
the assistant's own words), not only from a tool output that carries instruction-like text (solvi.perturb's rules;
injections="any": any such tool output escalates the call), your policies (@guard.policy(tools, on_fail="deny" | "escalate"): ordinary solvi hard checks over the arguments and the facts your app gives;@guard.fnfor computations
they read) and, optionally, an authorizer — a decider's yes / no "does the conversation authorize this call?"
(guard.make_authorizer(decider), perturb=2,guard.calibrate_authorizer(examples, risk=0.10)= act_guard).- The outcome:
allow(solvi runs the registered function:d.result, ord.errorwhen it raised),denyor
escalate, with the reasons in words (d.reasons,d.message()for the model), the candidate call and the evidence
(where each grounded argument is quoted). A failed deny check wins over a failed escalate check; an abstention (a fact
not given, an unsure authorizer) is an escalation.guard.resolve(d, approve, reviewer)records a person's answer and
makes an approved call.guard.session(context, facts)follows a conversation and feeds tool outputs back into it. - Each tool is a solvi System with one question,
verdict: every decision is a full response — trace, audit, stored with
meta["guard"](outcome, reasons, executed, the result's hash or the error) in a TraceStorage;guard.replay(id),
guard.replay_all(); the same call in the same conversation gives the same trace.guard.check/acheckdecide
without running anything;acallawaits async tools and policies. - Adapters (each imports its framework only when used):
solvi.agents.pydantic_ai.GuardedToolset(a WrapperToolset:
deny → ModelRetry, escalate → ApprovalRequired and deferred approval),solvi.agents.langgraph.guarded_tool_node(a
ToolNode with wrap_tool_call: deny → an error ToolMessage, escalate → interrupt / Command(resume=...)),
solvi.agents.openai_agents.guard_tools(a tool input guardrail + needs_approval: deny → reject_content, escalate →
an interruption to approve). Tested with pydantic-ai 2.51, langgraph 1.2.12 and openai-agents 0.22.3 and their
scripted models (dependency groupagents; the tests skip without them). solvi serve --guard catalog.py:guard --upstream CMD [--facts JSON] [--escalate elicit|deny] [--store]: an MCP proxy
in front of an MCP server —tools/listshows the declared tools (their schemas adopted from the server), every
tools/callpasses the guard; an escalation asks the user through MCP elicitation when the client supports it.solvi checklints a Guard (every tool's checks). examples/19_agent_guard.py: an
accounts-payable agent, scripted, through every case.
Agent guard: after a benchmark run (preview)
A run on AgentDojo (97 agent tasks, five kinds of prompt injection, gpt-oss-120b and Qwen3-235B) showed where the
default guard costs honest work. Every change below is opt-in, except the detector's, and keeps the provenance
guarantee of the default. Each has tests (tests/test_agents_next.py).
-
Middle mode:
Guard(tool_values="escalate")/tool(..., tool_values="escalate"). A user-only argument whose
value is not in the user's words but is in a tool output escalates (the new checkarguments_from_user, with the
quote and, in a tainted context, the instruction) instead of being denied.- What it relaxes: such a call is decided by a person instead of refused. Nothing is allowed on its own that the
default denies. A value found nowhere, or only in the assistant's or system's words, is still denied. The
escalation is never covered by a standing approval (policy_onlyis False). - Security cost: the guarantee for these values moves to the reviewer. In the run, a call with the attacker's
value reached the (simulated, strict) reviewer in 24–43% of attacked runs. Attacks that succeeded went from 2.1% /
1.6% to 1.9% / 2.7%: e-mails to real meeting participants carrying an attacker's link were approved. - Utility: with a reviewer, honest tasks solved rose by 7 and 16 points over the default with the same reviewer.
Without one, by nothing.
- What it relaxes: such a call is decided by a person instead of refused. Nothing is allowed on its own that the
-
URL matcher:
ground={"url": "url"}and"url_prefix". URLs are compared by parsing, not as tokens:- equal host (lower case, IDNA, no trailing dot, one leading
www.ignored), port (80 / 443 default), path (trailing
/ignored), query and fragment; the scheme is never downgraded (a writtenhttps://matches only an
https://call; a writtenhttp://or no scheme matches either); - never a match: userinfo (
good.com@evil.com), a host that only contains the name (evil.com/good.com,
good.com.evil.com), a backslash, other schemes,./..segments, a look-alike IDN; "url_prefix"lets the path continue a written one at a/(for reads only);solvi.agents.same_url/url_partsfor your own policies.
What it relaxes: a missing scheme or an http → https upgrade,
www., a trailing slash, a default port and the letter case of the host, and a
Unicode host equals its punycode form. Withurl_prefix, any sub-path of a written URL. In the run, web page reads
refused because the model addedhttp://went from 27 to 0. - equal host (lower case, IDNA, no trailing dot, one leading
-
guard.require_request(tools, intent, phrases=None, on_fail="escalate")is a policy for actions with no
user-given value (book, create an event, read a URL a document names). The call goes ahead only when the user's own
messages ask for this kind of action (solvi.agents.INTENTS: reserve, event, visit, pay, send, delete, invite,
post, share — English and Russian — or your own regular expressions); otherwise it escalates or is denied. It checks
the kind of action, not the call: a user who asked for any calendar event "asked" for one with an attacker's title,
which is the one attack that still passed. -
The detector sees more commands (the guard's rules only; a decider's
perturb=kis unchanged):- at the start of a sentence or after a colon: "Make a reservation for …", "…, and make a reservation", "Book … for
/ at …", "Visit / go to ", "Create … event / meeting / reminder"; - in Russian: "забронируй", "сделай бронирование", "зайди / перейди на сайт …", "создай событие …";
- "reserve" and "visit" join the verbs of "please … / you must …";
- a JSON or
reprtool output is read again with its escaped\nas line breaks. Before, every start-of-line rule
missed an instruction inside such an output.
False flags on honest text: 1.2% → 2.0% of AgentDojo's environment tex...
- at the start of a sentence or after a colon: "Make a reservation for …", "…, and make a reservation", "Book … for
solvi 0.6.1 — deterministic hashes of failed steps
- A failed step's value (MISSING) hashed as
repr(object()), which carries a memory address, so a trace with a failed step
hashed differently in every process and could not be replayed or verified from a store in another process. It now
hashes as{"missing": true}. Hashes of failed steps change once; nothing else changes.
solvi 0.6.0 — serving, catalog lint, several models, async, measured costs
Async execution: aask
await system.aask(state, names=None, order=None, store=True, timeout=None, speculate=False)next toask:
async defcatalog parts (fn, extract, check, rule, alternative producers) are awaited; steps run concurrently as
soon as the steps they read have finished; sync parts run inline, or in a worker thread (asyncio.to_thread) when
declaredblocking=True.- Early exit: by default in the phases of
ask(hard checks and what they read first), so no call starts thatask
would not make;speculate=Truestarts every ready step at once and cancels the pending calls that a failed hard
check makes unnecessary. Cancellingaaskcancels every pending call. - Timeouts:
timeout=(seconds) on a part (@cat.fn(timeout=2), extract, check, rule), per call (aask(timeout=)) or
for the System (System(timeout=)). A call that does not finish fails with "timed out after 2 s"; the questions that
need it abstain with guardtimeout— a new safeguard inres.safeguards, the audit,system.stats["timeouts"]and
safeguard_report()(listed once it fires) — and a producer that times out is followed by the next one. Replay does
not re-run a step that timed out. - The trace is the one
askwrites: records in flow order, the same answers and hashes whatever finished first — tested
on all 117 gallery cases and on examples 01, 03, 04, 09, 12 and 16, phased and speculative, and with storage,
concurrent asks, batched decisions andCascade/Vote/Route. ask, replay andfacts_forstill work on catalogs withasync defparts (each call awaited in an event loop of its
own).System.is_async(solvi.runtime.async_parts(catalog)) says whether a catalog has parts thataaskawaits;
solvi serveanswers such a System withaask(async HTTP endpoints, concurrent asks; the MCP server too).solvi.runtime.aexecuteis the async executor;executeandaexecuteshare one plan of phases.
Costs from measurements
System(..., producers="equivalent", costs="measured"): the cost-optimal planner (ModelStrategist(producers= "equivalent");producers="equivalent"on the System is now a shortcut for it) plans with the run times
system.costsmeasures instead of declared costs. Warm-up: a producer counts its measured time aftermin_samples
runs; before that its declaredcost=, or 0 ms when undeclared, so each is tried and measured. When the producer in use
slows down, the next plans switch; a producer unused forrecheckasks gets one more trial. Settings:
solvi.learned.MeasuredCosts(min_samples=3, recheck=50, alpha=None)(alpha: the smoothing ofsystem.costs).system.freeze_costs()fixes the planner's costs at what was measured (the choice stops changing; measuring goes on),
system.unfreeze_costs()resumes.- The plan record says why each path was chosen:
extra["costs"]lists, per fact with several usable producers, each
producer's cost and its source (measured,declared,warm-up,recheck,frozen: ...) and awhyline. ModelStrategist.plan(..., costs={producer: cost})takes costs from the caller (under the strategist's owncosts=).
solvi serve: HTTP, MCP and System One
solvi serve module:attr(orfile.py:attr) serves a System's questions over HTTP (solvi[serve]: FastAPI, uvicorn):
POST /ask(state in;Response.to_dict()out withstored_idandtrace_hash),POST /ask/{question},
GET /questions,GET /health. The OpenAPI document comes from the same pydantic types: each question's input state
schema (the given facts its flow reads, typed bySystem(inputs=...)or by their typed readers, the ones it cannot be
answered without as required;solvi.serve.question_inputs) and each response's answers as closed sets. The state is
not validated by the web layer: a wrong-typed field is rejected by solvi as usual (the answers that need it abstain,
safeguardtype_rejected).--store PATHsaves every answer with its trace to a TraceStorage.solvi serve --mcp: an MCP server over stdio, each question a tool whose input schema is the question's input state
schema; a call returns the answer, confidence, status, why and safeguards with the stored id. Uses the officialmcp
SDK (2.x,solvi[mcp]) when installed, else a built-in JSON-RPC server (initialize, ping, tools/list, tools/call).POST /v1/systemonebacked by a solvi decider (--decider path_or_hf_id,--model-name): the System One protocol
(choice → probabilities, noul → P(yes), score → expected level index with its legend), so solvi answers where a Jev /
Kev client points;solvi.systemoneround-trips against it.solvi serve --decider Xalone serves only this endpoint.System.response_schemais built bysolvi.schema.response_model(system, names=None)(the pydantic class);
solvi.strategist.given_facts(catalog)lists the facts a catalog reads and no part produces.
solvi check: catalog lint
solvi check module:attr(solvi.check.lint(system)): catalog lint with exit status 0 (no errors) / 1 / 2 (usage),
--strict(warnings fail),--json. Errors: a hard check whosethen=question never runs it (not read by the rule,
not incheckpoints: a failing check would be ignored),then=naming no question or an invalid answer, cycles,
questions no input can answer, producer / consumer andinputs=type conflicts, constraints that cannot hold (alone or
together; brute force over finite answer domains), constraints reading non-questions. Warnings: unused parts,then=on
soft checks, rules reading question names, disagreeing reader types, options the constraints always rule out, raising
constraints, and silent defaults —x or <literal>/.get(k, <literal>)in functions that read the input (# solvi: okaccepts one).
Several models, one decision
solvi.multi.Cascade([small, large]): ask the decision parts in order, answer with the first that does not escalate,
escalate when all do. The next model is asked only when needed;costs=[45, 137]reports the expected cost.solvi.multi.Vote([a, b], rule="all" | "majority"): answer when the rule holds and every agreeing part is sure;
disagreement escalates with the proposals listed. Parts of one model share a forward pass when they can.solvi.multi.Route({predicate or fact name: part}, default=part): code picks the part per input; only its model runs.- A combination is used wherever a decision part is (
cat.fn,.question(cat),System.teachteaches every part);
the parts must answer the same question (checked at construction); combinations nest. act_guard(examples, risk=0.10)on the combination: one threshold on every part's signal, chosen by conformal risk
control on the loss monotonized from above (a cascade's loss is not monotone in the threshold), so P(answered alone
and wrong) ≤ risk holds for the whole. Measured with solvi-base → solvi-large at risk 0.10: the risk stayed ≤ 10% on
every data set; the cascade answered 96% of ContractNLI alone at 64 ms per question against the large model's 97% at
137 ms; voting lowered the error among automatic answers on JSON questions from 2.1% to 0.4%. Alsoconformal.- The trace records every proposal (
extra["stages"]/["answered_by"],["votes"],["route"]/["routed"]) and
the models called (extra["calls"]); the audit lists each stage, vote or route; replay re-runs every stage and compares
the proposals, or — trusted or unavailable models — checks that the answer follows from the recorded proposals. examples/18_several_models.py: cascade, vote and route under one guarantee, with keyword stand-ins.
solvi 0.5.1 — escalation with a guarantee, any System One model, a release gate, stored decisions
Escalation with a guarantee
Measured on the 0.5.0 deciders: the shipped act threshold for "10% error" let through answers that were wrong 32–39% of
the time on typed-decisions and Taskmaster-2 (it holds on ContractNLI and JSON questions). The thresholds below keep their
promise on inputs like your calibration examples.
part.act_guard(examples, risk=0.10): conformal risk control on a few hundred labelled examples of your stream —
P(answered alone and wrong) ≤ risk, as a share of all questions. Measured on solvi-large with 300 examples: the risk stays
at 9.6–10.0% on every data set (typed-decisions answers 32% alone, ContractNLI 97%, JSON questions 99.6%). The result
also says how much must escalate at least when the model is often wrong (must_escalate_at_least).part.calibrate_for(examples, error=..., method="ltt"): learn-then-test — the error among the answers given alone ≤
error with probability ≥ 1 − delta; stricter, it often lets nothing through.method="empirical"is the 0.5.0 behaviour.part.conformal(examples, coverage=0.9): every decision carriesextra["candidates"], the answers that cannot be
ruled out; an escalation's message lists them for the person who takes over.- Every decision records what its threshold promises; the audit shows a
guaranteeline per answer, or says that there
is none because the thresholds were not calibrated on your data. solvi.calibration:crc_threshold,ltt_threshold,conformal_quantile,set_scores.
Safeguards
- Changed default: choice and multi-label decisions ask the model with the options in sorted order
(option_order="canonical"), so how a caller lists them cannot change the answer. On an independent stress test
(decision-models-under-pressure, 64 options) reordering the options changed 41% of solvi-large's answers in the given
order and 0.5% in the canonical one, at about the same accuracy. Options, probabilities and multi-label answers are still
shown in the caller's order.option_order="given"restores 0.5.0 (and its fingerprints);"average"averages over
rotations of the list. Parts whose options were not already sorted get a new fingerprint. min_margin=0.1: escalate a near tie between the two most probable answers (where a misleading text flips a choice).- An answer head with a NaN or infinite feature abstains instead of answering with confidence NaN (found by fuzzing).
- Quotes proposed by a model are shown in the audit as "in the text; support not checked" (the text match is checked;
whether the quote supports the answer is not).
Any System One model as a decider
solvi.systemone.systemone(base_url, model, api_key=None): a decider overPOST /v1/systemone— Jev and open servers
(Kev, Von, Laya-serve, Intern-Decision, …). Questions about one input go in one request; everything built on a decider
works: act_guard, conformal, fit / teach, audit, trace (which records the endpoint and model name).
Release gate and decision tests
- Honesty suite (
solvi.honesty,solvi honesty SET --baseline B): abstaining, "not stated", act vs escalate and traps on
a labelled set; three numbers — confident errors, coverage at 10% risk, share of quotes that back the answer (a proxy) —
and a non-zero exit when any gets worse. Run in CI and before publishing a model (docs/honesty.md). solvi test PATHand a pytest plugin: decision regression tests fromcases.json(the gallery format) — expected
answers, statuses and safeguards per case, trace replay,--fuzz Ninput mutations (docs/testing.md).
Models
- The deciders are now
solvi-ai/solvi-largeandsolvi-ai/solvi-base(the olddecide-large/decide-baseids
redirect).
Storage
TraceStorage(solvi.storage): stored responses with their whole traces —save,get(id),query(question=, answer=, status=, safeguard=, model=, since=, until=),iter,corrections,replay_all(system). Backends
JSONLStorage(append-only, one record per line) andSQLiteStorage(stdlib sqlite3, indexed; several writers). A hash
chain across stored records:verify()catches an edited, deleted, inserted or reordered record and a cut-off tail
(the stored head;verify(anchor=head)against a head kept elsewhere).quarantine(fact, value)lists the stored
decisions whose answers rest on a fact;forget(fact, value)reports what removing a given fact would touch (nothing is
deleted).System(..., storage=...)saves every ask (res.stored_id) and everyteach;ask(..., store=False)skips one.
journal="file.jsonl"is now aJSONLStorage: the 0.5 line keys are kept (plus the whole response and the chain
fields), 0.5 lines already in the file are kept and reported aslegacy;teachlines store dates as ISO strings.- Catalog fingerprint:
System.fingerprint()andtrace.fingerprint(the catalog's, the questions' and every flow
part's fingerprint: declarations, declared types and the code's syntax tree with the constants and same-module helpers
it reads;solvi.provenance.catalog_fingerprint).Trace.replaysays whether the catalog changed since the trace was
recorded and which parts;TraceStorage.query(catalog=fp). solvi.diff.diff(store, system): re-run stored decisions with a new catalog or model and list the answers, statuses,
safeguards and confidences that change, each with the first step that differs and why.Shadow(current, candidate, storage=...): answer with the current system, store the candidate's response and the differences.- A
solvicommand (alsopython -m solvi):solvi verify,solvi replay,solvi diffover a store. Result.whyshows set-valued facts in a fixed order (it depended onPYTHONHASHSEED), so stored responses hash the same
in every process.
solvi 0.5.0 — typed facts, typed decisions, answer primitives
Types declare questions, the model proposes, checks decide. Type hints on catalog functions are now fact types
(pydantic), the fields of a pydantic model are the questions a decider answers, and every answer — from a rule or a model —
can be "not stated", a span of the text, a ranking, a number with an interval, or carry evidence quotes, each checked
against the text. A code strategist plans around dead ends and picks the cheapest verified plan.
pip install -U solvi · guide ·
playground (new presets) · models:
solvi-ai/decide-large and
solvi-ai/decide-base (CPU / ONNX)
Known limits: the deciders are previews — read their cards: ahead of GLiNER2.5-Decide and Laya on typed questions over JSON
states, behind GLiNER2.5-Decide on zero-shot choice questions (level with it after fit on ~64 examples); the act / escalate
thresholds shipped with a model are indicative — calibrate on your own data before trusting them. The strategist's segment model and solvi.aliases are experimental and their weights are not published.
Type hints on catalog functions are the types of the facts (pydantic v2); untyped catalogs behave and hash exactly as
before. The same types declare the questions a decider model answers: types declare questions, the model proposes, checks
decide.
Models: the decider checkpoints published with this release are previews; each model card on
huggingface.co/solvi-ai has its measured numbers and limits. The strategist's model
weights are not published.
Typing
- Typed facts (
solvi.typed):def risk_score(risk_points: dict[str, float]) -> float— the catalog records each fact's
type (cat.types,cat.readers,flow.types) and checks every producer's return type against every consumer's argument
type when a part is registered; a definite mismatch raisesFactTypeErrornaming both functions (conservative:int→
float,str→date,dict→ model,X | None→Xpass). A typed rule's return type is checked against its
question's options when theSystemis built. - Run time: a typed part's arguments (given or computed) and its output (a Quote's / Decision's value) are validated and
coerced with pydanticTypeAdapters (cached per type; exact-type fast path; values that already passed the same type in
the run are not re-validated). A failure is rejected like an ungrounded quote — the fact is missing, the next producer
runs or dependent answers abstain — and is a new safeguard,type_rejected("type rejected"): in the step's error, the
audit,res.safeguardsandSystem.stats. A Literal / Enum return type is a closed set (outside it:outside_options);
an Enum answer is returned as its value.validategets the coerced value; replay re-runs the validation. - Answer types from Python types:
Answer.from_type(bool | Literal[...] | Enum | list[Literal[...]], ordinal=False);
Question(name, text)withoutanswer=takes it from its rule's return type. - Typed input state:
system.ask(model_instance)(a pydanticBaseModel: its fields are the given facts);
System(..., inputs=Model)validates dict requests — fields with defaults become given facts, a field that fails is left
out and reported (res.trace.rejected, safeguardtype_rejected). - Serialization (
solvi.schema, pydantic models):model_dump(mode),to_json(),model_validate(data, catalog=),
from_json(text, catalog=),model_json_schema()onResponse,Result,Trace,Record,Question,AnswerType;
system.response_schema()has each answer as its closed set. Withcatalog=(or the System), typed values that JSON
cannot carry (dates, enums, models) are restored from the facts' types, so a loaded trace replays with the same hashes. - pydantic (
>=2) is a core dependency; it is imported only for typed parts, BaseModel inputs and serialization
(import solvidoes not load it; pydantic ships with Pyodide, so the browser playground can use it). - Faster asks: a value read by several steps is hashed once per run (gallery runners up to 14% faster).
- Example 14 (typed customs desk); gallery 10 (procurement) retrofitted with pydantic documents and typed functions (same
answers; the audit shows the given documents as models). - Trace note: records of typed parts hash their coerced values; untyped traces are unchanged.
- Hand-written extractors: a Quote without its own
sourcepoints into the extractor's text —docif the function
reads it, else its only argument, else its onlystr-typed argument; an ambiguous signature raises at registration and
asks for the new@cat.extract(source="...").
Typed decisions
- The decider (
solvi.decide) answers typed questions; L14b–L14e checkpoints (l14b_decider v1) load, score and hash
exactly as before.- Question kinds from types:
choice(Literal[...], an Enum; with "other" as an abstain threshold),multi
(list[Literal[...]]),score(solvi.typed.Scale[Literal[...]], 2–10 ordered levels → an ordinal answer; the value
is the median, the expected level is recorded),noul(bool→ the value True / False, answered yes / no).
model.decision(name, task, fact, Scale[...])(ortype=,kind=),model.decisions(PydanticModel, fact)(one part
per field: its type the kind, its description the task),model.questions(cat, PydanticModel, fact);
Answer.from_type(Scale[...])is ordinal;solvi.typed.question_kind,Scale,Ordinal. - Input: a text, or a state — a dict, list, pydantic model or dataclass — serialized by
solvi.decide.state_textas key
paths (customer.tier: pro), exactly the L14f training serialization ("paths"; also "tree" and "json", as the checkpoint
declares). A decision reading several facts serializes{fact: value}. - Output per question: probabilities, a calibrated confidence (a temperature per kind) and act / escalate. The model's act
signal (an act head, optionally through a shipped act calibrator) below its threshold rejects the decision as the new
safeguard model escalated (guard="escalated",system.stats["model_escalated"], the audit); without one,
escalate_below=escalates by calibrated confidence as low confidence.act_threshold=,target_error=(the
checkpoint's threshold for an error rate),use_act=False;part.calibrate_for(examples, error=0.05)picks the
threshold for a target error rate. An escalated decision's answer abstains saying what it would have answered; a
fallback producer runs if there is one. Provenance staysdecided;record.extrahas the act probability. - Several questions per forward pass: when the checkpoint declares
multi_question, the strategist groups decision
parts reading the same facts with the same model (flow.batches) and the executor scores each group in one pass
(model.passescounts them; the block layout of L14f — input encoded once, questions do not see each other — with a
fallback to one question per pass); records name their shared pass (extra["pass"]) and replay re-scores it.
model.decide_pass(input, parts). Catalogs without decisions do no extra work. adapt/fit/teachper kind: a free shift per option (choice, multi), an ordinal-aware tilt and spread over the
levels (score), one yes−no bias (noul);System.teachmaps answers to the decision's labels (True→ yes).- Checkpoint capabilities in
solvi_decide.json(formatsl14b_decider v1,l14f typed v1,solvi_decide v2): modes,
markers, head columns, noul labels, state serialization, multi-question layout, temperatures per kind, thresholds, act
head (column, temperature, calibrator, thresholds per target error) — the contract is
docs/decide_format.md.DecideModel.load(..., multi_question=, act=)overrides them for
experiments;model.caps. - Records, flows and their JSON carry the new
extra/batches(records without them hash as before).
- Question kinds from types:
Answer primitives
- Answer primitives — every answer is a value and a confidence, declared by types, from plain rules, learned parts and
model decisions alike (solvi.primitives; guide: "Answer primitives"; examples/16_primitives.py):- "Not stated":
solvi.Unknown(typeNotStated;Maybe[T]=T | NotStated;Answer.maybe(t)) is a real answer —
the text does not state it — with a confidence, distinct from "no" and from an abstention (None).result.not_stated,
res.not_stated,res.overall["not_stated"]; constraints seeUnknownand joint decoding can choose it; it
round-trips through JSON ("not_stated": true, probability key"<not stated>"). - Evidence:
Claim(value, evidence=[Quote | str], confidence=, source=)from any part,Decision(..., evidence=)from
a model; strings are located in the text, every quote must be literally in its given text at its offsets — else the
output is rejected (safeguard "grounding rejected": the fact is missing, the next producer runs, else the answer
abstains). Recorded inrecord.extra["evidence"](hashed, replayed, tampering caught),result.evidence, shown in the
audit and counted in the support (quoted/quoted_by_model).Question(require_evidence=True): an answer without a
quote abstains — the new safeguard evidence missing (guard="evidence_missing",system.stats["evidence_missing"],
listed bysafeguard_report()once it fires). - Span:
Span[T]/Answer.span(source=, type=)— an exact substring of a given text (a Quote, or a text that is
located), always grounded, coerced toTwith pydantic (a failure: "type rejected");result.span. - *...
- "Not stated":