Skip to content

Releases: solvi-ai/solvi

solvi 1.0.1

Choose a tag to compare

@mxkuzn mxkuzn released this 04 Oct 06:42

Fixed

  • A compact record now keeps the answer a hard check's then= function gave, and replay checks it. In 1.0.0 a
    compact record (record="compact" or the compact part of "sample:N") left out the trace record of kind then
    (name then:<question>) that a then= function of facts writes, so replay_all reported every decision such a
    function answered as not_kept ("not verified") instead of checking it — 51 of 100 compact decisions in the
    playground's compact-journal demo. Compact records now keep it whole, like a model step (its value, error and the
    hashes of the facts it read), and replay_all / store.rederive re-run the function on the re-computed facts and
    compare: an edited value, or a function that now gives another answer, is a mismatch at then:<question>. A kept
    record that does not match its step hash in the record is reported as integrity (it was not_kept). A then record
    costs about 0.4 KB in a compact record (the demo: 1.8 → 2.0 KB a decision); benchmarks/journal_size.py does not
    change (its workloads have no then= functions). Full records do not change. Compact records written by 1.0.0
    still load, verify and replay as before: they do not hold the then record, so those decisions stay not_kept
    (tests/fixtures/store_1_0_0).

solvi 1.0.0 — two levels, knowledge and agents

Choose a tag to compare

@mxkuzn mxkuzn released this 04 Oct 05:20

1.0 makes solvi two levels. The high level is ready systems you configure: solvi.build (decisions from
labelled examples with a promise on the errors), solvi.Agent (new: acting in an environment), solvi.Guard (an
agent's tool calls) and solvi.Knowledge (new: what a system knows and from whom — facts with their sources, exact
retraction, an agenda of goals and gates, an action model learned from outcomes). The low level, solvi.core, is
the building blocks they are made of, with every extension point exported, documented and covered by a conformance
check. What works but has not shown a measured gain is moved to solvi.experimental, marked in every decision
that uses it, with a deadline (1.2) to graduate or go. Field reports from building on solvi brought quotes matched on
a normalized view, quotes from several labelled sources, then= that wires its hard check and can compute the answer,
res.checks, a compact journal, budgets for generation and refinement, and outcomes as labels. Python 3.11 or newer.
Every 0.9 import path keeps working through 1.0.x with a warning, and solvi migrate rewrites your code. What was
tried for this release and did not meet its bar is listed at the end.

Breaking changes and migration

If your code imports only the 23 names solvi still exports (below), it runs unchanged. Every other 0.9 import path
still works in 1.0.x: it imports the very same module and warns (SolviDeprecationWarning) with the path to use;
1.1 removes the old paths. Run solvi migrate PATH once (below), then your tests with
-W error::solvi.SolviDeprecationWarning. What breaks now, without a warning period:

  • Python 3.11 or newer (see Changed for the fingerprints this fixes).
  • The removed modules (table under Removed) raise ModuleNotFoundError.
  • The methods moved off the classes (table below) raise an AttributeError that names the new call.
  • then= wires its hard check. A hard check whose then names a question now runs in that question's flow by
    itself; requires= is no longer needed for it, and other questions' flows do not change. Before, forgetting
    requires let the question be answered as if the check had passed. A strategist of your own that leaves such a check
    out is refused when the System is built; solvi check (then_not_in_flow) keeps reporting it for a strategist
    swapped in afterwards. A system that relied on the check not running for that question now gets the forced answer
    when it fails.

The new layout: solvi's modules moved into packages that say what you can rely on — solvi and solvi.solutions
are the ready-made systems, solvi.core and its areas (solvi.core.types, solvi.core.runtime, ...) are the
building blocks they are made of, and solvi.experimental holds what may still change.

  • solvi migrate PATH rewrites your code (.py and .md files) to the new paths: imports, from solvi import storage, dotted paths in strings such as monkeypatch.setattr("solvi.llm.urlopen", ...) or
    "solvi.hooks:rules_system". solvi migrate PATH --check changes nothing and exits 1 when a file would change.
  • Stored decisions, calibration files and fingerprints do not change: a fingerprint records the 0.9 module of a moved
    class or function (the table solvi._deprecate.MOVED), so decisions stored by 0.7–0.9 replay, and a store written by
    1.0 is read by 0.9 tools the same way. A stored module:qualname that names a 0.9 module loads without a warning.
  • What solvi exports: 23 names (__all__) — the entry points build (solvi.solutions.decisions.build, was
    solvi.auto.build), Agent (new, solvi.solutions.agent), Guard (solvi.solutions.guard.Guard, was
    solvi.agents.Guard) and Knowledge (new, solvi.solutions.knowledge), Budget, and the shared vocabulary
    Catalog, Question, Answer, System, Response, Quote, Claim, Decision, Fail, Unknown, Span,
    Maybe, Rank, Estimate, Scale, Bins, SolviDeprecationWarning, ExperimentalWarning. The other 0.9 names (JSONLStorage and the other stores,
    TraceStorage, Trace, Record, Result, MISSING, Shadow, AnswerType, NotStated, FactTypeError) are
    imported from their modules (table below); from solvi import JSONLStorage works in 1.0.x with a warning.
  • solvi.auto.AutoSystem is now solvi.solutions.decisions.DecisionSystem (what solvi.build returns; the old name
    works in 1.0.x with a warning). build(slow=...) no longer compiles a written specification itself (that would make
    it import the experimental compiler): compile it first (solvi.experimental.compile.compile_spec) and pass the
    result; writer= and inputs= raise a TypeError saying so.

Where each module went

you imported (0.9) import now (1.0) level
solvi.agents solvi.solutions.guard high level: ready to use
solvi.agents.confirm solvi.solutions.guard.confirm high level: ready to use
solvi.agents.guard solvi.solutions.guard high level: ready to use
solvi.agents.intents solvi.solutions.guard.intents high level: ready to use
solvi.agents.mcp solvi.experimental.mcp experimental: may change; removed in 1.2 unless it graduates
solvi.agree solvi.core.slow.agree low level: building blocks
solvi.audit solvi.core.store.audit low level: building blocks
solvi.auto solvi.solutions.decisions high level: ready to use
solvi.calibfile solvi.core.calibfile low level: building blocks
solvi.calibration solvi.core.calibration low level: building blocks
solvi.charts solvi.experimental.charts experimental: may change; removed in 1.2 unless it graduates
solvi.charts.check solvi.experimental.charts.check experimental: may change; removed in 1.2 unless it graduates
solvi.charts.propose solvi.experimental.charts.propose experimental: may change; removed in 1.2 unless it graduates
solvi.charts.render solvi.experimental.charts.render experimental: may change; removed in 1.2 unless it graduates
solvi.charts.spec solvi.experimental.charts.spec experimental: may change; removed in 1.2 unless it graduates
solvi.command solvi._command internal
solvi.compile solvi.experimental.compile experimental: may change; removed in 1.2 unless it graduates
solvi.costs solvi.core.costs low level: building blocks
solvi.counterfactual solvi.experimental.counterfactual experimental: may change; removed in 1.2 unless it graduates
solvi.decide solvi.core.deciders low level: building blocks
solvi.decide.adapt solvi.core.deciders.adapt low level: building blocks
solvi.decide.backends solvi.core.deciders.backends low level: building blocks
solvi.decide.capabilities solvi.core.deciders.capabilities low level: building blocks
solvi.decide.gate solvi.core.deciders.gate low level: building blocks
solvi.decide.kinds solvi.core.deciders.kinds low level: building blocks
solvi.decide.model solvi.core.deciders.model low level: building blocks
solvi.decide.part solvi.core.deciders.part low level: building blocks
solvi.decide.state solvi.core.deciders.state low level: building blocks
solvi.decide.wire solvi.core.deciders.wire low level: building blocks
solvi.diff solvi.core.store.diff low level: building blocks
solvi.dispatch solvi.core.dispatch low level: building blocks
solvi.drift solvi.core.guarantees.drift low level: building blocks
solvi.episode solvi.core.knowledge.episodes low level: building blocks
solvi.extract_long solvi.core.extract low level: building blocks
solvi.generate solvi.core.slow.generate low level: building blocks
solvi.guarantee solvi.core.guarantees.guarantee low level: building blocks
solvi.heads solvi.core.deciders.heads low level: building blocks
solvi.honesty solvi.testing.honesty high level: ready to use
solvi.hooks solvi.experimental.hooks experimental: may change; removed in 1.2 unless it graduates
solvi.i18n solvi.core._i18n internal
solvi.inputs solvi.core._inputs internal
solvi.learning solvi.experimental.learning experimental: may change; removed in 1.2 unless it graduates
solvi.llm solvi.core.deciders.llm low level: building blocks
solvi.loader solvi._loader internal
solvi.longdoc solvi.core.deciders.longdoc low level: building blocks
solvi.lora solvi.experimental.lora experimental: may change; removed in 1.2 unless it graduates
solvi.memory solvi.core.knowledge.memory low level: building blocks
solvi.multi solvi.core.deciders.combine low level: building blocks
solvi.openset solvi.core.guarantees.openset low level: building blocks
solvi.perturb solvi.core.deciders.perturb low level: building blocks
solvi.primitives solvi.core.primitives low level: building blocks
solvi.provenance solvi.core.provenance low level: building blocks
solvi.refine solvi.core.slow.refine low level: building blocks
solvi.remote solvi.core.deciders._remote internal
solvi.report solvi.core.store.report low level: building blocks
solvi.rulelist solvi.core.deciders.rulelist low level: building blocks
solvi.runtime solvi.core.runtime low level: building blocks
solvi.sandbox solvi.experimental.compile.sandbox experimental: may change; removed in 1.2 unless it graduates
solvi.scaffold solvi.cli._scaffold internal
solvi.schema solvi.core.schema low level: building blocks
solvi.search `solvi.cor...
Read more

solvi 0.9.0 — System 1 and System 2

Choose a tag to compare

@mxkuzn mxkuzn released this 03 Oct 13:40

0.9 is about two ways of deciding in one system, in Kahneman's sense: System 1, fast and cheap — rules, checks, a
fitted head, a model under a guarantee — answers when it is sure; System 2, slow and deliberate — an LLM, a re-ask
loop, a search — is woken when System 1 is unsure or surprised, and a person gets what neither can answer. A
dispatcher puts both in one recorded, replayable decision within a budget and can be calibrated on the inputs System 1
hands over; a system report tells the owner who answered, at what cost, and whether the promise held; System 2 can also
write what System 1 then runs fast — a policy text compiled into catalog parts, with a person settling what the drafts
dispute, and into an agent guard. Search over candidates runs about twice as fast, the nine-task benchmark stand reruns
in CI, and a showcase on the world map of Pokémon Red puts the pieces together. The names 0.8 renamed and kept with a
warning are removed. What was tried for this release and did not meet its bar is listed at the end.

Breaking changes

The old names that 0.8 kept working with a SolviDeprecationWarning are gone. An old keyword now raises a TypeError
and an old attribute or method an AttributeError; both say "X was renamed in 0.8 and removed in 0.9: use Y". The five
old modules are gone (importing one raises ModuleNotFoundError). If your code ran under 0.8 without a
SolviDeprecationWarning (pytest -W error::solvi.SolviDeprecationWarning finds them all), it runs under 0.9
unchanged. solvi.SolviDeprecationWarning itself stays, for later renames.

Removed (old → what to use):

removed use
module solvi.fast solvi.heads
module solvi.learned solvi.costs (CostBook, MeasuredCosts) and solvi.strategist (OrderModel, ProducerPolicy, Binary, scalar_row)
module solvi.rules solvi.rulelist
module solvi.strategy_model solvi.segment_model
module solvi.extract_model (SpanExtractor) solvi.extract_long.LongSpanExtractor
System(inputs=), system.inputs System(input_model=), system.input_model
System(costs=), system.costs System(cost_policy=), system.cost_book
System(journal=path) System(storage=JSONLStorage(path)) or storage="file.jsonl"
ask(names=), aask(names=), Service.ask(names=) / aask(names=), Shadow.ask(names=) questions=
Question(checkpoints=), q.checkpoints, part.question(cat, checkpoints=) (also on a combination) requires=, q.requires
res.computed_state, res.computed_state_text(lang) res.state_text(lang)
res.textin res.read
system.teach(source=), store.save_correction(source=) label_source=
system.learn_rule(facts=) features=
system.calibrate(question, states, truth) system.calibrate(question, [(state, answer), ...])
system.safeguard_report() system.safeguard_summary()
shadow.report() shadow.summary()
system.fit_fast(...) system.fit(..., select=False)
System.learning(harvest_rules=) (ignored in 0.8) — (teach rule outcomes with label_source="rule")
trace.value(name) res.values[name]; a given fact: trace.init[name]
answer_type.rank(v) answer_type.options.index(v)
model.decision(escalate_below=, act_threshold=, target_error=, unknown=) and the same in decide, adapt, fit, teach, reset, decisions, questions min_confidence=, min_act=, max_error=, not_stated=
a pydantic field's json_schema_extra keys "escalate_below", "act_threshold", "target_error" (now a ValueError) "min_confidence", "min_act", "max_error"
model.has_unknown model.has_not_stated
model.long_len, part.long_len max_len_long
act_guard(risk=) (a part and a combination), adapt_lora(risk=), CorrectionMemory.calibrate(risk=), Guard.calibrate_authorizer(risk=) max_risk=
calibrate_for(error=) (a part and a combination) max_error=
the result keys "coverage" and "target_error" of calibrate_for "answered", "max_error"
the result key "calls" of a combination's act_guard; combination.usage() and its key "per_question" "calls_per_question"; combination.calls()
a combination's decide(x=), teach(x=), adapt(inputs=) text=, text=, texts=
FastHead.update(row, answer), Binary.observe(row, y) teach(...)
FastHead.cv_acc, Head.loo_acc (each head's number under the other's name) FastHead.loo_acc, Head.cv_acc
ModelStrategist() without a model CostStrategist() (ModelStrategist now needs its model)
CostStrategist(fallback=, fallbacks=) / ModelStrategist(...), .fallback, .fallbacks on_failure=, keep_alternatives=
MultiSpanExtractor.fit(docs, spans), predict_doc(text) fit([(text, spans), ...]), predict(text[, field])
store.forget(fact, value) store.where_is(fact, value)
JSONLStorage(path, catalog=) (every backend), store.get(id, catalog=) system=
store.query(catalog=fingerprint) query(catalog_fp=)
solvi.testing.check(system, case, state) solvi.testing.run_case(...)
a honesty case's "gold" (now a ValueError); a solvi test case's "gold" (now reported as a problem of the case) "expected"
solvi hook ... --model M (exits 1 with what to do, so an old installed hook blocks nothing), $SOLVI_HOOK_MODEL --decider M (reinstall the hooks), $SOLVI_HOOK_DECIDER
solvi.serve.Guard (the ASGI middleware) solvi.serve.AccessGuard
agents.Guard(facts=[names]) Guard(fact_names=[names])
the agent adapters' declare= (guard_tool, guard_wrappers, guarded_tool_node, GuardedToolset) auto_declare=
run_proxy(context_messages=, context_chars=), solvi.agents.mcp.Proxy(...) the same max_messages=, max_chars=
solvi.llm._error_text (private) solvi.llm.error_text

Stores, calibration files and fingerprints are unchanged by these removals: the stored keys stay as they were (a
question still hashes its requires under the key checkpoints, a part still stores escalate_below). A test loads,
verifies and replays stores written by 0.8.0 and by 0.7.1 (tests/test_store_0_8_0.py, tests/test_store_0_7_1.py).

One fingerprint changes once: the code fingerprint of a catalog (its rules, functions and checks) no longer depends on
the Python version (3.12 and 3.13 print a function's syntax tree differently), so it differs from the one 0.8 computed
for the same code. Decisions stored by 0.8 still load, verify and replay; anything that compares a stored catalog
fingerprint with the current one (store.query(catalog_fp=...), a version check of your own) sees the catalog as
changed once. Model and question fingerprints, and a decision part's, are unchanged (tests/test_store_0_8_0.py).

New

  • One entry point: solvi.auto.build (preview). The question, labelled examples, the promise (max_risk= or
    max_error=) and, optionally, a slow path (an LLM, a decision part, a System, a function or a compiled
    specification), a budget and a store in; a ready System 1 and dispatcher out, with ask(), report() and
    explain() — a plain account of what System 1 is, the signal its guarantee reads, how the examples were split, the
    promise and its threshold, who answers each slice it hands over and what is not covered. It composes what the library
    has: the catalog's rule, your own fitted part (learner=) or a head fitted on the computed facts; the act
    probability, the confidence or the computed number that best separates right from wrong; an open-set gate for a
    choice among more than two options; Dispatcher.calibrate on examples System 1's guarantee did not see — only when
    System 1 actually hands the slow path a slice, else those examples calibrate System 1. Measured on four tasks of the
    stand against the hand-written setups: as good or better with every promise kept on eval, at a quarter to half of
    the code. Limits: it lands closer to the promised level than the hand-written setups (risk 9.3% of 10%, 0.9% of 1%;
    error 4.9% of 5% after a shift), it rarely gives the slow path anything (inputs unlike the examples go to a person,
    and on the hard slice System 1's own guess was often the better answer), and its drift flag came later than the
    hand-written one (191 vs 68 requests after a shift). Example: examples/24_one_entry_point.py.
  • Who answers: solvi.dispatch (experimental). Dispatcher(system1, SlowPath(system2), ...) asks System 1
    first; when its own signals say its answer cannot be given alone (below its guarantee, outside the open-set gate,
    an abstention, a broken constraint, low agreement), the slow path answers or checks — a System, a re-ask loop
    (propose=, solvi.refine) or a search (space=, solvi.search) — and when that cannot answer either, or the
    budget (Budget(usd=, calls=, ms=), per decision and in total) is spent, a person gets the input with both
    candidates and the reasons, never a guess. Drift and a sampled share (supervise=) make the slow path check System
    1's answers instead. Every decision is one hash-chained record with its cost (dollars from the recorded tokens),
    and d.replay(res) / d.replay_all() re-check it without calling a model. Guide: "Who answers".
  • Calibrating who answers on the hard slice: Dispatcher.calibrate(examples, max_risk=...) measures, on each
    slice of the inputs System 1 hands over, System 1's own would-be answer, the slow path's, and the slow path's when it
    agrees with System 1, and picks one answerer per slice with a threshold — or a person — so that all answers given
    alone keep the promise. The inputs handed over are the hard ones, where an LLM that is rarely wrong on average can be
    wrong often and System 1's own guess can be the better answer; calibrate measures that instead of assum...
Read more

solvi 0.8.0 — one name per concept, any model first

Choose a tag to compare

@mxkuzn mxkuzn released this 02 Oct 22:03

0.8 gives every concept one name, makes the surface smaller and the decider protocol one, puts any model first (an LLM
through solvi.llm, a decision service, or a local checkpoint for offline use), and fixes what an independent audit of
0.7 found. Most renamed names still work in 0.8 with a warning that names the new one; they go in 0.9. The warning is
solvi.SolviDeprecationWarning, a FutureWarning, so Python shows it to you (once per old name per process) in your
own code too, not only in tests and __main__ as it would a DeprecationWarning.
Stores, calibration files and fingerprints written by 0.7.1 load, verify and replay unchanged (a test replays stores
written by 0.7.1).

Breaking changes

Removed or changed without a working alias:

  • Every System option after questions is keyword-only: System(cat, qs, "file.jsonl") is a TypeError —
    System(cat, qs, storage="file.jsonl"). Every ask / aask option after the questions is keyword-only too:
    ask(state, ["q"], 4) → ask(state, ["q"], workers=4).
  • act_guard (of a part and of a Cascade / Vote / Route), calibrate_for, adapt_lora, Guard.calibrate_authorizer:
    every option after the examples is keyword-only — act_guard(examples, 0.1) → act_guard(examples, max_risk=0.1).
    systemone(url, model, key, 30.0) → systemone(url, model, key, timeout=30.0).
  • The octonion signature left the package: sign(obj, "octonion"), reading an octonion signature and solvi verify --sign --alg octonion raise a ValueError pointing to benchmarks/octonion_signature.py (the default syndrome code
    locates the same changes, 4x smaller, ~60x faster). solvi verify has no --alg option.
  • model.decision(...) raises for an option the question's kind does not use, where it accepted and ignored it (some
    changed the part's fingerprint): k= off a ranking, score_value= off a score question, bins= / unit= /
    coverage= off a number question, other= on a score or yes/no question, min_margin= on a multi-label one, top_k=
    / rerank= without long=, min_act= / max_error= on a checkpoint without an act head, kind= that contradicts
    multi=True. decisions(schema, fields=[...]) raises for a field the schema does not have.
  • A wrong key, model or URL for a System One service (HTTP 401, 403, 404, another 4xx except 400 / 413 / 422) raises
    SystemOneError, as solvi.llm raises LLMError (both are solvi.remote.RemoteError); it used to escalate every
    decision, and solvi models check then reported "answered alone 0.0%" and exited 0.
  • scorer.usage of an LLM decider and extra["llm"]["usage"] count input_tokens / output_tokens /
    reasoning_tokens, as a System One decider does (they were prompt_tokens / completion_tokens).
  • guarantee["signal"] of a Cascade / Vote / Route is a name — "shared", "shared-rank" — as a part's is ("act",
    "confidence"); it was a sentence. Fingerprints and calibration files of 0.7 stay valid.
  • Escalation texts of remote models: "the LLM server refused the request: HTTP 400 — ..." (was "invalid input for the
    endpoint: ..."), "the LLM server did not answer after 3 attempts: ..." (was "no answer from ... after 3 attempts").
  • Command lines that pass an option their mode does not read are refused (status 2) instead of being ignored: solvi serve in HTTP, --mcp and --guard modes; solvi ask --report with --json, --audit or --lang; --backend /
    --api-key without --decider; solvi report --id with period filters. POST /ask and POST /ask_text refuse a
    body key they do not know (422) — a misspelled "question" used to ask every question.
  • solvi ask --text --json prints the text read as "read" (was "textin"), as POST /ask_text does.
  • Removed, nothing called them: Catalog.producer, strategy.fact_of, Selection.expanded, serve.RequestTimeout,
    SystemOneScorer.question, ModelStrategist(max_expand=) / search(max_expand=).
  • New in this release and renamed before it is published (no alias): solvi.many.fit → fits (its record key too),
    DriftMonitor.calibrate → set_reference, EpisodeView.revisits(kind, key) → revisits(key, kind) with the key
    required, Episode.note(kind, key, value) → note(kind, key), episode.Chooser(escalate_below=) →
    min_confidence=; System.guarantee, solvi.guarantee.calibrate and OpenSetGate.calibrate take max_risk= /
    max_error= only.

Renamed, the old name works in 0.8 with a warning (old → new). Every old name warns once per process with
solvi.SolviDeprecationWarning (a FutureWarning, shown by Python's default filters): "X is deprecated since 0.8 and
will be removed in 0.9: use Y". To find them all at once, run your tests with -W error::solvi.SolviDeprecationWarning;
to silence them, warnings.filterwarnings("ignore", category=solvi.SolviDeprecationWarning).

0.7 0.8
System(inputs=Model), system.inputs System(input_model=Model), system.input_model
System(costs="measured"), system.costs System(cost_policy="measured"), system.cost_book
System(journal=path) System(storage=JSONLStorage(path)) (or storage="file.jsonl")
system.ask(state, names=...), aask(names=), Service.ask(names=), Shadow.ask(names=) questions=
Question(checkpoints=[...]), q.checkpoints, part.question(cat, checkpoints=) requires=, q.requires
res.computed_state, res.computed_state_text(lang) res.state_text(lang)
res.textin res.read
system.teach(..., source=), store.save_correction(..., source=) label_source=
system.learn_rule(question, examples, facts=) features=
system.calibrate(question, states, truth) system.calibrate(question, [(state, answer), ...])
system.safeguard_report(), shadow.report() safeguard_summary(), summary()
Trace.value(name) res.values[name] (a given fact: trace.init[name])
AnswerType.rank(v) answer_type.options.index(v)
model.decision(escalate_below=, act_threshold=, target_error=, unknown=) (and decide, decisions, questions, a field's json_schema_extra) min_confidence=, min_act=, max_error=, not_stated=
model.has_unknown, model.long_len, part.long_len has_not_stated, max_len_long
act_guard(risk=), adapt_lora(risk=), CorrectionMemory.calibrate(risk=), Guard.calibrate_authorizer(risk=) max_risk=
calibrate_for(error=); its result's "coverage", "target_error" max_error=; "answered", "max_error"
a combination's act_guard()["calls"], combination.usage() ("per_question") "calls_per_question", combination.calls()
FastHead.update(row, answer), Binary.observe(row, y) teach(...)
FastHead.cv_acc, Head.loo_acc (each head's number under the other's name) FastHead.loo_acc, Head.cv_acc
ModelStrategist() without a model; fallback=, fallbacks= CostStrategist(); on_failure=, keep_alternatives=
solvi.fast, solvi.learned, solvi.rules, solvi.strategy_model solvi.heads; solvi.costs + solvi.strategist; solvi.rulelist; solvi.segment_model
solvi.extract_model.SpanExtractor solvi.extract_long.LongSpanExtractor
MultiSpanExtractor.fit(docs, spans), predict_doc(text) fit([(text, spans), ...]), predict(text[, field])
store.forget(fact, value) (it never deleted anything) store.where_is(fact, value)
JSONLStorage(path, catalog=) (all backends), store.get(id, catalog), store.query(catalog=fingerprint) system=; query(catalog_fp=)
solvi.testing.check(system, case, state) run_case(...)
a honesty case's "gold" "expected", as in solvi test cases
solvi hook ... --model M, $SOLVI_HOOK_MODEL (hooks installed by 0.7 keep working) --decider M, $SOLVI_HOOK_DECIDER
solvi.serve.Guard (the ASGI middleware) solvi.serve.AccessGuard
agents.Guard(facts=[names]) Guard(fact_names=[names])
the adapters' declare=True auto_declare=True
run_proxy(context_messages=, context_chars=) max_messages=, max_chars=
System.learning(harvest_rules=True) — (it harvested nothing in the usual wiring; ignored)

Structure: solvi.decide is a package of layers (kinds, state, wire, capabilities, backends, adapt,
gate, part, model) and re-exports every name it had; the combinations' base is public as
solvi.multi.Combination; ExperimentalWarning and the "not stated" option name live in solvi.core; the helpers
every command shares are in solvi.command, so no library module imports from solvi.cli; every module declares its
public names in __all__, and the API reference shows only those (a helper that is not in __all__ may change without
notice).

New

  • Any model first. The README and the guide present the decider as whichever model you have: an LLM through
    solvi.llm (the core install is enough), a System One service, or a local checkpoint such as solvi-base for offline
    or cheap use — with solvi-base's model-card numbers where it is offered.
  • solvi.remote: the client every remote model shares (solvi.llm, solvi.systemone, solvi.generate, the chart
    proposer) — the endpoint, the key in the header only, retries with backoff, token counts under one set of names, and
    one policy for HTTP errors: a wrong key, model or URL raises, a refused input escalates that decision, no answer is
    retried and then escalates.
  • One decider protocol. A decision part and a Cascade / Vote / Route have the same public methods, with the same
    signatures and result keys: act_guard(examples, *, max_risk, signal, groups, min_group, delta) (a combination's
    result with calls_per_question, cost, scale and answered_by besides), and now on a combination too
    calibrate_for(examples, *, max_error, signal, method, delta) (one shared threshold for...
Read more

solvi 0.7.1 — solvi behind a coding agent's hooks (preview), gallery for coding agents, System One hardened, benchmark vs LLMs

Choose a tag to compare

@mxkuzn mxkuzn released this 29 Sep 10:01

solvi behind a coding agent's hooks (preview)

  • solvi hook pre-edit --rules rules.toml: a PreToolUse hook for Claude Code's Edit, Write and MultiEdit. It reads the
    proposed change from the hook's JSON, works out the added lines with their line numbers in the file after the edit,
    and asks a small solvi System one question, edit ∈ {allow, deny, ask}, whose hard checks are the rules whose path
    globs match: forbid (regular expressions over the added lines), require (over the file after the edit),
    forbid_calls and require_def (Python, from the parsed code), a rule with no checks (any change to these paths), and
    a fuzzy question a decider answers when an added line matches when. It answers "deny" with the rule, the lines and
    the rule's reason (the agent reads it and can fix the change), "ask" (the user confirms), or nothing (Claude Code's own
    permissions apply; --approve answers an explicit allow). A fuzzy rule blocks only with a calibration file
    (act_guard: P(answered alone and wrong) ≤ risk on labelled changes); without one its "yes" asks, and without a model
    a triggered question asks. A check that cannot run, instruction-like text addressed to a reviewer in the added lines,
    and a hook that fails all ask — never a silent allow.
  • solvi hook pick-skill --skills-dir .claude/skills: a UserPromptSubmit hook that picks one skill from the skills'
    names and descriptions (the words a prompt shares with each, weighted by rarity; or a decider's choice with --model)
    and adds one line naming it as additionalContext; silent on "none", a near tie or a slash command.
  • --model: a local checkpoint (never downloaded), a System One service (systemone:URL#model; solvi serve --decider ... --model-name ... keeps a local model loaded), an OpenAI-compatible endpoint (llm:URL#model, key from the
    environment) or your own decider. Calibrate a rule's question with solvi calibrate solvi.hooks:rules_system RULE_answer labels.jsonl --risk 0.1 ($SOLVI_HOOK_RULES, $SOLVI_HOOK_MODEL) and name the file in the rule.
  • Every decision is stored with its trace (.solvi/traces/hooks.jsonl by default, any TraceStorage with --store),
    so solvi verify and solvi report work on it; parallel hooks share one chain (a file lock), and the store opens from
    its head, not by reading every record. solvi hook audit [ID] prints a stored decision's audit and replays it against
    the current rules.
  • solvi hook install merges the hook entries into the project's .claude/settings.json (other hooks and settings
    stay; its own are replaced, not doubled), writes the sample rules (solvi hook sample-rules: no secrets in source, no
    employee data taken from the browser in app/api, reversible migrations, no eval or shell strings, a person for CI
    workflows) and prints what it changed; solvi hook uninstall removes exactly its entries.
  • Codex (preview): --agent codex writes .codex/hooks.json, reads the apply_patch envelope, and answers in Codex's
    dialect (no "ask": a deny that says a person must confirm).
  • Speed: without a model a hook call is one short process — 105–121 ms on a laptop (median), with a store of 3000
    decisions.
  • examples/22_coding_agent_hooks.py: a session in a temporary project — a clean
    edit, a rule broken, a comment that tries to talk past the rules, two prompts, the verified store and one audit.

Benchmark: solvi vs asking an LLM

  • docs/vs_llm.md — the same inputs and written policies given to solvi, four LLMs and a hosted
    decision model, directly and inside solvi: refunds, 3-way invoice matching, the gallery and routing bank messages.
    Strong reasoning LLMs followed the rules nearly perfectly and solvi was not more accurate there; the differences are
    cost, latency, answers that do not change with option order, replay and hard checks. benchmarks/vs_llm/ has the
    data, the written policies, the runner (the solvi arm offline and free; any OpenAI-compatible or System One endpoint)
    and every raw answer, so bench.py score --check recomputes the published tables without an API key. The
    playground gains a "solvi vs LLM" tab on the same data.

Gallery: helpers for coding agents

Three new entries for decisions a coding agent (such as Claude Code or Codex) meets on every task. Each runs offline:
the deciders are keyword stand-ins calibrated with act_guard on synthetic, seeded examples, and each README shows
solvi-large, any OpenAI-compatible LLM (solvi.llm) or a System One service in front, with the keywords as the
fallback, and says what the offline rules cannot read. All three are playground presets.

  • gallery/13_pre_edit_rule_check: before the agent writes a file, the project's rules for that path (per glob) are
    checked → allow / block / escalate, and the rules broken, with their lines. Secrets, browser storage read in
    app/api/** and irreversible migrations (Python's own parser: a downgrade() that does something, RunPython /
    RunSQL with their reverse) block by hard checks; a CI workflow edit and a migration that does not parse go to a
    person. "Auth checks go through require_role()" and "no personal data in log lines" are a decider's question each —
    out of scope → the decider (act_guard, risk 5%; perturb=2) → a person, as fallback producers of one fact. A
    # reviewer: ignore the rules above comment is not read by the code checks and flips the decider, which escalates.
    Also: a pre-edit hook script and a Guard policy on a write_file tool. 16 cases.
  • gallery/14_review_triage: seven yes/no risk questions → quick or full review. Three are code (dependencies, a large
    or unfocused change, logic without tests), four a decider's (auth, public API, migrations or deletion, security),
    each calibrated at risk 2.5% so that P(a risky change goes to quick review) ≤ 10% (a union bound). On 2000 fresh
    synthetic changes: 0.2% risky and quick, 22% quick overall. The change generator (synthetic_changes(n, seed)) is in
    task.py; one known miss (a permission change without its usual words) is a case. 14 cases.
  • gallery/15_skill_picker: a prompt → exactly one of ten skills (with look-alike pairs), none, or a person. A
    confident "none", an abstention and a wrong pick are kept apart; a skill the user names is cited, one named inside an
    instruction-like passage of pasted text is not; a near tie (min_margin) escalates with the conformal candidates; a
    production deploy needs the user's own word "production" (a hard check). 16 cases.

Changes

  • import solvi no longer imports numpy: the answer heads load it on first use, and hashing a fingerprint only looks for
    arrays when numpy is already loaded. The pydantic models of solvi.schema build their validators on first use.
  • tomli is a dependency on Python 3.10 (rules files are TOML).
  • solvi.systemone: systemone(..., extra_body=...) (and SystemOneScorer(..., extra_body=...)) merges
    server-specific fields into every request — OpenRouter's provider routing, user — with the rules of
    solvi.llm's: copied, merged under solvi's own fields, model / state / questions refused with ValueError, part
    of the fingerprint (without it the fingerprint is unchanged).
  • solvi.systemone answers "not stated" (unknown=True, Maybe[...]): the question gets one more option, "not
    stated", with a description; a yes/no question that allows it is asked as a choice over yes / no / not stated. Its
    probability competes with the options' as with solvi.llm and the checkpoints with a "not stated" output: the
    decision is Unknown and the question built on it abstains. Before, decision(unknown=True) raised ValueError.
  • solvi.systemone answers multi-label questions: one noul per option in the same request, an option chosen at the
    model's multi threshold (0.5), the confidence the least sure option's max(p, 1 − p) — what act_guard calibrates on;
    with "not stated" allowed, one more noul for it. SystemOneScorer.questions(item) gives an item's questions;
    question(item) still gives the one question of a single-question item. Spans and evidence stay refused.
  • solvi.systemone records per decision, in extra["systemone"], the endpoint, the model name (served_by when the
    service names another), the request's ms and, when the service reports them, its usage tokens and cost — for
    the whole request (questions: how many questions it answered); the scorer sums them in usage and cost.
    costs="measured" already plans on each part's measured run time, the request included.
  • Behaviour change: solvi.systemone handles a service that fails as solvi.llm does, instead of raising:
    network errors, timeouts, a broken connection (IncompleteRead, a reset), 408 / 409 / 429 and 5xx are retried
    (retries=2, backoff=1.0 s doubling), then the decision escalates ("the System One service did not answer after 3
    attempts: ...") and is not cached, so the next ask tries again; another 4xx escalates at once with the service's error
    text, a gateway's wrapped cause included (OpenRouter's error.metadata.raw); a reply that breaks the contract (a
    missing answer or probability) escalates ("invalid System One output — ..."). The API key is never in the reason.
  • Behaviour change: a System One model is deterministic=False by default, as solvi.llm's: replay checks the
    recorded output instead of calling the service again. systemone(..., deterministic=True) keeps the old re-run for
    a local server whose output is reproducible.

Fixes

  • The audit's guarantee line said "none for some decisions: their thresholds were not calibrated" when two decisions
    behind one answer carried the same promise (identical promises were counted once against the number of decisions).
  • A decision escalated by min_margin (a near tie) is now a low confidence ...
Read more

solvi 0.7.0 — text in, agent guard (preview), several models with an LLM stage, learning from corrections, long documents, LoRA (experimental)

Choose a tag to compare

@mxkuzn mxkuzn released this 29 Sep 06:52

The agent guard (solvi.agents) ships as a preview: its hard line is provenance (a value found only in a tool's
output never grounds an argument that must come from the user) and your policies; detecting injected instructions in
text is a heuristic second line. System.learning is experimental and off unless you call it; part.adapt_lora is
experimental too. Three code reviews and three adversarial passes ran before this release; their fixes are listed under
"Fixes before release".

Text in: entry points

  • system.entry_points(names=None): the questions as entry points — name, text and the typed input fields each one reads
    (type, description, required), from the same schemas as solvi serve; ep.tool() is the function-calling form.
  • solvi.textin.TextIn(system, decider, extractor=None, ...): read(text) → a TextRead — the entry point the decider
    picks (a choice over the entry points and their descriptions; escalates below min_confidence=0.6, on a near tie
    min_margin=0.1 or on the decider's act signal), and each input field read by span extraction with a quote and a
    deterministic parser per type: numbers ("1,500.50", "1.5 million", "2k", "полтора миллиона"), dates ("2026-09-12",
    "12.09.2026", "12 September", "12 сентября"; year-less and relative dates only with today=), enums by label or
    synonym, booleans, strings (with patterns=). A field is read, not_stated, unparsed, unsure or unsupported;
    required fields not read are in read.missing, and read.clarify() asks for them — nothing is guessed.
  • Extractors: the decider's span pointer (DeciderExtractor, when the checkpoint has one) or CueExtractor (deterministic
    candidates of the field's type nearest after a cue word); any object with find(text, FieldSpec) → [Quote].
  • system.ask_text(text | TextRead, decider=None, *, textin=None, question=None) (and aask_text): TextIn + ask in one
    trace. The text is a given fact (request_text); the entry point (textin, provenance decided) and each field
    (textin:<field>, provenance quoted, with the extractor's fingerprint, the parser and its arguments) are hash-chained
    records. The audit shows the fields as quoted by a model — never given, not in the deterministic share — and an answer's
    confidence is at most the reading's. Replay re-checks each quote, re-parses it and checks the flow read that value. An
    escalated entry point runs nothing: the likely questions abstain with guard escalated. res.textin is the TextRead.
  • A dialogue: tin.update(read, next_message) reads the next turn over the whole dialogue and lists changes (old value,
    new value, quote); "not A-10457 but A-10475" changes the field to the new value.

Guarding an agent's tool calls

  • solvi.agents.Guard: an agent proposes a tool call ({"name", "arguments"} — data, never code; OpenAI, LangChain,
    Anthropic and MCP shapes are read by ToolCall.parse) and solvi checks it as a proposal: the tool is in the catalog
    (@guard.tool on typed functions, guard.declare(name, schema=...) for a pydantic model or a JSON schema), the
    arguments validate against its types (unknown arguments are errors), the ground= arguments are quoted from the
    conversation (strings literally, numbers as number tokens, lists item by item; ground_from= the roles allowed — never
    the assistant's own words), not only from a tool output that carries instruction-like text (solvi.perturb's rules;
    injections="any": any such tool output escalates the call), your policies (@guard.policy(tools, on_fail="deny" | "escalate"): ordinary solvi hard checks over the arguments and the facts your app gives; @guard.fn for computations
    they read) and, optionally, an authorizer — a decider's yes / no "does the conversation authorize this call?"
    (guard.make_authorizer(decider), perturb=2, guard.calibrate_authorizer(examples, risk=0.10) = act_guard).
  • The outcome: allow (solvi runs the registered function: d.result, or d.error when it raised), deny or
    escalate, with the reasons in words (d.reasons, d.message() for the model), the candidate call and the evidence
    (where each grounded argument is quoted). A failed deny check wins over a failed escalate check; an abstention (a fact
    not given, an unsure authorizer) is an escalation. guard.resolve(d, approve, reviewer) records a person's answer and
    makes an approved call. guard.session(context, facts) follows a conversation and feeds tool outputs back into it.
  • Each tool is a solvi System with one question, verdict: every decision is a full response — trace, audit, stored with
    meta["guard"] (outcome, reasons, executed, the result's hash or the error) in a TraceStorage; guard.replay(id),
    guard.replay_all(); the same call in the same conversation gives the same trace. guard.check / acheck decide
    without running anything; acall awaits async tools and policies.
  • Adapters (each imports its framework only when used): solvi.agents.pydantic_ai.GuardedToolset (a WrapperToolset:
    deny → ModelRetry, escalate → ApprovalRequired and deferred approval), solvi.agents.langgraph.guarded_tool_node (a
    ToolNode with wrap_tool_call: deny → an error ToolMessage, escalate → interrupt / Command(resume=...)),
    solvi.agents.openai_agents.guard_tools (a tool input guardrail + needs_approval: deny → reject_content, escalate →
    an interruption to approve). Tested with pydantic-ai 2.51, langgraph 1.2.12 and openai-agents 0.22.3 and their
    scripted models (dependency group agents; the tests skip without them).
  • solvi serve --guard catalog.py:guard --upstream CMD [--facts JSON] [--escalate elicit|deny] [--store]: an MCP proxy
    in front of an MCP server — tools/list shows the declared tools (their schemas adopted from the server), every
    tools/call passes the guard; an escalation asks the user through MCP elicitation when the client supports it.
  • solvi check lints a Guard (every tool's checks). examples/19_agent_guard.py: an
    accounts-payable agent, scripted, through every case.

Agent guard: after a benchmark run (preview)

A run on AgentDojo (97 agent tasks, five kinds of prompt injection, gpt-oss-120b and Qwen3-235B) showed where the
default guard costs honest work. Every change below is opt-in, except the detector's, and keeps the provenance
guarantee of the default. Each has tests (tests/test_agents_next.py).

  • Middle mode: Guard(tool_values="escalate") / tool(..., tool_values="escalate"). A user-only argument whose
    value is not in the user's words but is in a tool output escalates (the new check arguments_from_user, with the
    quote and, in a tainted context, the instruction) instead of being denied.

    • What it relaxes: such a call is decided by a person instead of refused. Nothing is allowed on its own that the
      default denies. A value found nowhere, or only in the assistant's or system's words, is still denied. The
      escalation is never covered by a standing approval (policy_only is False).
    • Security cost: the guarantee for these values moves to the reviewer. In the run, a call with the attacker's
      value reached the (simulated, strict) reviewer in 24–43% of attacked runs. Attacks that succeeded went from 2.1% /
      1.6% to 1.9% / 2.7%: e-mails to real meeting participants carrying an attacker's link were approved.
    • Utility: with a reviewer, honest tasks solved rose by 7 and 16 points over the default with the same reviewer.
      Without one, by nothing.
  • URL matcher: ground={"url": "url"} and "url_prefix". URLs are compared by parsing, not as tokens:

    • equal host (lower case, IDNA, no trailing dot, one leading www. ignored), port (80 / 443 default), path (trailing
      / ignored), query and fragment; the scheme is never downgraded (a written https:// matches only an
      https:// call; a written http:// or no scheme matches either);
    • never a match: userinfo (good.com@evil.com), a host that only contains the name (evil.com/good.com,
      good.com.evil.com), a backslash, other schemes, . / .. segments, a look-alike IDN;
    • "url_prefix" lets the path continue a written one at a / (for reads only);
    • solvi.agents.same_url / url_parts for your own policies.

    What it relaxes: a missing scheme or an http → https upgrade, www., a trailing slash, a default port and the letter case of the host, and a
    Unicode host equals its punycode form. With url_prefix, any sub-path of a written URL. In the run, web page reads
    refused because the model added http:// went from 27 to 0.

  • guard.require_request(tools, intent, phrases=None, on_fail="escalate") is a policy for actions with no
    user-given value (book, create an event, read a URL a document names). The call goes ahead only when the user's own
    messages ask for this kind of action (solvi.agents.INTENTS: reserve, event, visit, pay, send, delete, invite,
    post, share — English and Russian — or your own regular expressions); otherwise it escalates or is denied. It checks
    the kind of action, not the call: a user who asked for any calendar event "asked" for one with an attacker's title,
    which is the one attack that still passed.

  • The detector sees more commands (the guard's rules only; a decider's perturb=k is unchanged):

    • at the start of a sentence or after a colon: "Make a reservation for …", "…, and make a reservation", "Book … for
      / at …", "Visit / go to ", "Create … event / meeting / reminder";
    • in Russian: "забронируй", "сделай бронирование", "зайди / перейди на сайт …", "создай событие …";
    • "reserve" and "visit" join the verbs of "please … / you must …";
    • a JSON or repr tool output is read again with its escaped \n as line breaks. Before, every start-of-line rule
      missed an instruction inside such an output.

    False flags on honest text: 1.2% → 2.0% of AgentDojo's environment tex...

Read more

solvi 0.6.1 — deterministic hashes of failed steps

Choose a tag to compare

@mxkuzn mxkuzn released this 28 Sep 14:29
  • A failed step's value (MISSING) hashed as repr(object()), which carries a memory address, so a trace with a failed step
    hashed differently in every process and could not be replayed or verified from a store in another process. It now
    hashes as {"missing": true}. Hashes of failed steps change once; nothing else changes.

solvi 0.6.0 — serving, catalog lint, several models, async, measured costs

Choose a tag to compare

@mxkuzn mxkuzn released this 28 Sep 13:49

Async execution: aask

  • await system.aask(state, names=None, order=None, store=True, timeout=None, speculate=False) next to ask:
    async def catalog parts (fn, extract, check, rule, alternative producers) are awaited; steps run concurrently as
    soon as the steps they read have finished; sync parts run inline, or in a worker thread (asyncio.to_thread) when
    declared blocking=True.
  • Early exit: by default in the phases of ask (hard checks and what they read first), so no call starts that ask
    would not make; speculate=True starts every ready step at once and cancels the pending calls that a failed hard
    check makes unnecessary. Cancelling aask cancels every pending call.
  • Timeouts: timeout= (seconds) on a part (@cat.fn(timeout=2), extract, check, rule), per call (aask(timeout=)) or
    for the System (System(timeout=)). A call that does not finish fails with "timed out after 2 s"; the questions that
    need it abstain with guard timeout — a new safeguard in res.safeguards, the audit, system.stats["timeouts"] and
    safeguard_report() (listed once it fires) — and a producer that times out is followed by the next one. Replay does
    not re-run a step that timed out.
  • The trace is the one ask writes: records in flow order, the same answers and hashes whatever finished first — tested
    on all 117 gallery cases and on examples 01, 03, 04, 09, 12 and 16, phased and speculative, and with storage,
    concurrent asks, batched decisions and Cascade / Vote / Route.
  • ask, replay and facts_for still work on catalogs with async def parts (each call awaited in an event loop of its
    own). System.is_async (solvi.runtime.async_parts(catalog)) says whether a catalog has parts that aask awaits;
    solvi serve answers such a System with aask (async HTTP endpoints, concurrent asks; the MCP server too).
  • solvi.runtime.aexecute is the async executor; execute and aexecute share one plan of phases.

Costs from measurements

  • System(..., producers="equivalent", costs="measured"): the cost-optimal planner (ModelStrategist(producers= "equivalent"); producers="equivalent" on the System is now a shortcut for it) plans with the run times
    system.costs measures instead of declared costs. Warm-up: a producer counts its measured time after min_samples
    runs; before that its declared cost=, or 0 ms when undeclared, so each is tried and measured. When the producer in use
    slows down, the next plans switch; a producer unused for recheck asks gets one more trial. Settings:
    solvi.learned.MeasuredCosts(min_samples=3, recheck=50, alpha=None) (alpha: the smoothing of system.costs).
  • system.freeze_costs() fixes the planner's costs at what was measured (the choice stops changing; measuring goes on),
    system.unfreeze_costs() resumes.
  • The plan record says why each path was chosen: extra["costs"] lists, per fact with several usable producers, each
    producer's cost and its source (measured, declared, warm-up, recheck, frozen: ...) and a why line.
  • ModelStrategist.plan(..., costs={producer: cost}) takes costs from the caller (under the strategist's own costs=).

solvi serve: HTTP, MCP and System One

  • solvi serve module:attr (or file.py:attr) serves a System's questions over HTTP (solvi[serve]: FastAPI, uvicorn):
    POST /ask (state in; Response.to_dict() out with stored_id and trace_hash), POST /ask/{question},
    GET /questions, GET /health. The OpenAPI document comes from the same pydantic types: each question's input state
    schema (the given facts its flow reads, typed by System(inputs=...) or by their typed readers, the ones it cannot be
    answered without as required; solvi.serve.question_inputs) and each response's answers as closed sets. The state is
    not validated by the web layer: a wrong-typed field is rejected by solvi as usual (the answers that need it abstain,
    safeguard type_rejected). --store PATH saves every answer with its trace to a TraceStorage.
  • solvi serve --mcp: an MCP server over stdio, each question a tool whose input schema is the question's input state
    schema; a call returns the answer, confidence, status, why and safeguards with the stored id. Uses the official mcp
    SDK (2.x, solvi[mcp]) when installed, else a built-in JSON-RPC server (initialize, ping, tools/list, tools/call).
  • POST /v1/systemone backed by a solvi decider (--decider path_or_hf_id, --model-name): the System One protocol
    (choice → probabilities, noul → P(yes), score → expected level index with its legend), so solvi answers where a Jev /
    Kev client points; solvi.systemone round-trips against it. solvi serve --decider X alone serves only this endpoint.
  • System.response_schema is built by solvi.schema.response_model(system, names=None) (the pydantic class);
    solvi.strategist.given_facts(catalog) lists the facts a catalog reads and no part produces.

solvi check: catalog lint

  • solvi check module:attr (solvi.check.lint(system)): catalog lint with exit status 0 (no errors) / 1 / 2 (usage),
    --strict (warnings fail), --json. Errors: a hard check whose then= question never runs it (not read by the rule,
    not in checkpoints: a failing check would be ignored), then= naming no question or an invalid answer, cycles,
    questions no input can answer, producer / consumer and inputs= type conflicts, constraints that cannot hold (alone or
    together; brute force over finite answer domains), constraints reading non-questions. Warnings: unused parts, then= on
    soft checks, rules reading question names, disagreeing reader types, options the constraints always rule out, raising
    constraints, and silent defaults — x or <literal> / .get(k, <literal>) in functions that read the input (# solvi: ok accepts one).

Several models, one decision

  • solvi.multi.Cascade([small, large]): ask the decision parts in order, answer with the first that does not escalate,
    escalate when all do. The next model is asked only when needed; costs=[45, 137] reports the expected cost.
  • solvi.multi.Vote([a, b], rule="all" | "majority"): answer when the rule holds and every agreeing part is sure;
    disagreement escalates with the proposals listed. Parts of one model share a forward pass when they can.
  • solvi.multi.Route({predicate or fact name: part}, default=part): code picks the part per input; only its model runs.
  • A combination is used wherever a decision part is (cat.fn, .question(cat), System.teach teaches every part);
    the parts must answer the same question (checked at construction); combinations nest.
  • act_guard(examples, risk=0.10) on the combination: one threshold on every part's signal, chosen by conformal risk
    control on the loss monotonized from above (a cascade's loss is not monotone in the threshold), so P(answered alone
    and wrong) ≤ risk holds for the whole. Measured with solvi-base → solvi-large at risk 0.10: the risk stayed ≤ 10% on
    every data set; the cascade answered 96% of ContractNLI alone at 64 ms per question against the large model's 97% at
    137 ms; voting lowered the error among automatic answers on JSON questions from 2.1% to 0.4%. Also conformal.
  • The trace records every proposal (extra["stages"] / ["answered_by"], ["votes"], ["route"] / ["routed"]) and
    the models called (extra["calls"]); the audit lists each stage, vote or route; replay re-runs every stage and compares
    the proposals, or — trusted or unavailable models — checks that the answer follows from the recorded proposals.
  • examples/18_several_models.py: cascade, vote and route under one guarantee, with keyword stand-ins.

solvi 0.5.1 — escalation with a guarantee, any System One model, a release gate, stored decisions

Choose a tag to compare

@mxkuzn mxkuzn released this 28 Sep 12:49

Escalation with a guarantee

Measured on the 0.5.0 deciders: the shipped act threshold for "10% error" let through answers that were wrong 32–39% of
the time on typed-decisions and Taskmaster-2 (it holds on ContractNLI and JSON questions). The thresholds below keep their
promise on inputs like your calibration examples.

  • part.act_guard(examples, risk=0.10): conformal risk control on a few hundred labelled examples of your stream —
    P(answered alone and wrong) ≤ risk, as a share of all questions. Measured on solvi-large with 300 examples: the risk stays
    at 9.6–10.0% on every data set (typed-decisions answers 32% alone, ContractNLI 97%, JSON questions 99.6%). The result
    also says how much must escalate at least when the model is often wrong (must_escalate_at_least).
  • part.calibrate_for(examples, error=..., method="ltt"): learn-then-test — the error among the answers given alone ≤
    error with probability ≥ 1 − delta; stricter, it often lets nothing through. method="empirical" is the 0.5.0 behaviour.
  • part.conformal(examples, coverage=0.9): every decision carries extra["candidates"], the answers that cannot be
    ruled out; an escalation's message lists them for the person who takes over.
  • Every decision records what its threshold promises; the audit shows a guarantee line per answer, or says that there
    is none because the thresholds were not calibrated on your data.
  • solvi.calibration: crc_threshold, ltt_threshold, conformal_quantile, set_scores.

Safeguards

  • Changed default: choice and multi-label decisions ask the model with the options in sorted order
    (option_order="canonical"), so how a caller lists them cannot change the answer. On an independent stress test
    (decision-models-under-pressure, 64 options) reordering the options changed 41% of solvi-large's answers in the given
    order and 0.5% in the canonical one, at about the same accuracy. Options, probabilities and multi-label answers are still
    shown in the caller's order. option_order="given" restores 0.5.0 (and its fingerprints); "average" averages over
    rotations of the list. Parts whose options were not already sorted get a new fingerprint.
  • min_margin=0.1: escalate a near tie between the two most probable answers (where a misleading text flips a choice).
  • An answer head with a NaN or infinite feature abstains instead of answering with confidence NaN (found by fuzzing).
  • Quotes proposed by a model are shown in the audit as "in the text; support not checked" (the text match is checked;
    whether the quote supports the answer is not).

Any System One model as a decider

  • solvi.systemone.systemone(base_url, model, api_key=None): a decider over POST /v1/systemone — Jev and open servers
    (Kev, Von, Laya-serve, Intern-Decision, …). Questions about one input go in one request; everything built on a decider
    works: act_guard, conformal, fit / teach, audit, trace (which records the endpoint and model name).

Release gate and decision tests

  • Honesty suite (solvi.honesty, solvi honesty SET --baseline B): abstaining, "not stated", act vs escalate and traps on
    a labelled set; three numbers — confident errors, coverage at 10% risk, share of quotes that back the answer (a proxy) —
    and a non-zero exit when any gets worse. Run in CI and before publishing a model (docs/honesty.md).
  • solvi test PATH and a pytest plugin: decision regression tests from cases.json (the gallery format) — expected
    answers, statuses and safeguards per case, trace replay, --fuzz N input mutations (docs/testing.md).

Models

  • The deciders are now solvi-ai/solvi-large and solvi-ai/solvi-base (the old decide-large / decide-base ids
    redirect).

Storage

  • TraceStorage (solvi.storage): stored responses with their whole traces — save, get(id), query(question=, answer=, status=, safeguard=, model=, since=, until=), iter, corrections, replay_all(system). Backends
    JSONLStorage (append-only, one record per line) and SQLiteStorage (stdlib sqlite3, indexed; several writers). A hash
    chain across stored records: verify() catches an edited, deleted, inserted or reordered record and a cut-off tail
    (the stored head; verify(anchor=head) against a head kept elsewhere). quarantine(fact, value) lists the stored
    decisions whose answers rest on a fact; forget(fact, value) reports what removing a given fact would touch (nothing is
    deleted).
  • System(..., storage=...) saves every ask (res.stored_id) and every teach; ask(..., store=False) skips one.
    journal="file.jsonl" is now a JSONLStorage: the 0.5 line keys are kept (plus the whole response and the chain
    fields), 0.5 lines already in the file are kept and reported as legacy; teach lines store dates as ISO strings.
  • Catalog fingerprint: System.fingerprint() and trace.fingerprint (the catalog's, the questions' and every flow
    part's fingerprint: declarations, declared types and the code's syntax tree with the constants and same-module helpers
    it reads; solvi.provenance.catalog_fingerprint). Trace.replay says whether the catalog changed since the trace was
    recorded and which parts; TraceStorage.query(catalog=fp).
  • solvi.diff.diff(store, system): re-run stored decisions with a new catalog or model and list the answers, statuses,
    safeguards and confidences that change, each with the first step that differs and why. Shadow(current, candidate, storage=...): answer with the current system, store the candidate's response and the differences.
  • A solvi command (also python -m solvi): solvi verify, solvi replay, solvi diff over a store.
  • Result.why shows set-valued facts in a fixed order (it depended on PYTHONHASHSEED), so stored responses hash the same
    in every process.

solvi 0.5.0 — typed facts, typed decisions, answer primitives

Choose a tag to compare

@mxkuzn mxkuzn released this 28 Sep 06:25

Types declare questions, the model proposes, checks decide. Type hints on catalog functions are now fact types
(pydantic), the fields of a pydantic model are the questions a decider answers, and every answer — from a rule or a model —
can be "not stated", a span of the text, a ranking, a number with an interval, or carry evidence quotes, each checked
against the text. A code strategist plans around dead ends and picks the cheapest verified plan.

pip install -U solvi · guide ·
playground (new presets) · models:
solvi-ai/decide-large and
solvi-ai/decide-base (CPU / ONNX)

Known limits: the deciders are previews — read their cards: ahead of GLiNER2.5-Decide and Laya on typed questions over JSON
states, behind GLiNER2.5-Decide on zero-shot choice questions (level with it after fit on ~64 examples); the act / escalate
thresholds shipped with a model are indicative — calibrate on your own data before trusting them. The strategist's segment model and solvi.aliases are experimental and their weights are not published.


Type hints on catalog functions are the types of the facts (pydantic v2); untyped catalogs behave and hash exactly as
before. The same types declare the questions a decider model answers: types declare questions, the model proposes, checks
decide.

Models: the decider checkpoints published with this release are previews; each model card on
huggingface.co/solvi-ai has its measured numbers and limits. The strategist's model
weights are not published.

Typing

  • Typed facts (solvi.typed): def risk_score(risk_points: dict[str, float]) -> float — the catalog records each fact's
    type (cat.types, cat.readers, flow.types) and checks every producer's return type against every consumer's argument
    type when a part is registered; a definite mismatch raises FactTypeError naming both functions (conservative: int →
    float, str → date, dict → model, X | None → X pass). A typed rule's return type is checked against its
    question's options when the System is built.
  • Run time: a typed part's arguments (given or computed) and its output (a Quote's / Decision's value) are validated and
    coerced with pydantic TypeAdapters (cached per type; exact-type fast path; values that already passed the same type in
    the run are not re-validated). A failure is rejected like an ungrounded quote — the fact is missing, the next producer
    runs or dependent answers abstain — and is a new safeguard, type_rejected ("type rejected"): in the step's error, the
    audit, res.safeguards and System.stats. A Literal / Enum return type is a closed set (outside it: outside_options);
    an Enum answer is returned as its value. validate gets the coerced value; replay re-runs the validation.
  • Answer types from Python types: Answer.from_type(bool | Literal[...] | Enum | list[Literal[...]], ordinal=False);
    Question(name, text) without answer= takes it from its rule's return type.
  • Typed input state: system.ask(model_instance) (a pydantic BaseModel: its fields are the given facts);
    System(..., inputs=Model) validates dict requests — fields with defaults become given facts, a field that fails is left
    out and reported (res.trace.rejected, safeguard type_rejected).
  • Serialization (solvi.schema, pydantic models): model_dump(mode), to_json(), model_validate(data, catalog=),
    from_json(text, catalog=), model_json_schema() on Response, Result, Trace, Record, Question, AnswerType;
    system.response_schema() has each answer as its closed set. With catalog= (or the System), typed values that JSON
    cannot carry (dates, enums, models) are restored from the facts' types, so a loaded trace replays with the same hashes.
  • pydantic (>=2) is a core dependency; it is imported only for typed parts, BaseModel inputs and serialization
    (import solvi does not load it; pydantic ships with Pyodide, so the browser playground can use it).
  • Faster asks: a value read by several steps is hashed once per run (gallery runners up to 14% faster).
  • Example 14 (typed customs desk); gallery 10 (procurement) retrofitted with pydantic documents and typed functions (same
    answers; the audit shows the given documents as models).
  • Trace note: records of typed parts hash their coerced values; untyped traces are unchanged.
  • Hand-written extractors: a Quote without its own source points into the extractor's text — doc if the function
    reads it, else its only argument, else its only str-typed argument; an ambiguous signature raises at registration and
    asks for the new @cat.extract(source="...").

Typed decisions

  • The decider (solvi.decide) answers typed questions; L14b–L14e checkpoints (l14b_decider v1) load, score and hash
    exactly as before.
    • Question kinds from types: choice (Literal[...], an Enum; with "other" as an abstain threshold), multi
      (list[Literal[...]]), score (solvi.typed.Scale[Literal[...]], 2–10 ordered levels → an ordinal answer; the value
      is the median, the expected level is recorded), noul (bool → the value True / False, answered yes / no).
      model.decision(name, task, fact, Scale[...]) (or type=, kind=), model.decisions(PydanticModel, fact) (one part
      per field: its type the kind, its description the task), model.questions(cat, PydanticModel, fact);
      Answer.from_type(Scale[...]) is ordinal; solvi.typed.question_kind, Scale, Ordinal.
    • Input: a text, or a state — a dict, list, pydantic model or dataclass — serialized by solvi.decide.state_text as key
      paths (customer.tier: pro), exactly the L14f training serialization ("paths"; also "tree" and "json", as the checkpoint
      declares). A decision reading several facts serializes {fact: value}.
    • Output per question: probabilities, a calibrated confidence (a temperature per kind) and act / escalate. The model's act
      signal (an act head, optionally through a shipped act calibrator) below its threshold rejects the decision as the new
      safeguard model escalated (guard="escalated", system.stats["model_escalated"], the audit); without one,
      escalate_below= escalates by calibrated confidence as low confidence. act_threshold=, target_error= (the
      checkpoint's threshold for an error rate), use_act=False; part.calibrate_for(examples, error=0.05) picks the
      threshold for a target error rate. An escalated decision's answer abstains saying what it would have answered; a
      fallback producer runs if there is one. Provenance stays decided; record.extra has the act probability.
    • Several questions per forward pass: when the checkpoint declares multi_question, the strategist groups decision
      parts reading the same facts with the same model (flow.batches) and the executor scores each group in one pass
      (model.passes counts them; the block layout of L14f — input encoded once, questions do not see each other — with a
      fallback to one question per pass); records name their shared pass (extra["pass"]) and replay re-scores it.
      model.decide_pass(input, parts). Catalogs without decisions do no extra work.
    • adapt / fit / teach per kind: a free shift per option (choice, multi), an ordinal-aware tilt and spread over the
      levels (score), one yes−no bias (noul); System.teach maps answers to the decision's labels (True → yes).
    • Checkpoint capabilities in solvi_decide.json (formats l14b_decider v1, l14f typed v1, solvi_decide v2): modes,
      markers, head columns, noul labels, state serialization, multi-question layout, temperatures per kind, thresholds, act
      head (column, temperature, calibrator, thresholds per target error) — the contract is
      docs/decide_format.md. DecideModel.load(..., multi_question=, act=) overrides them for
      experiments; model.caps.
    • Records, flows and their JSON carry the new extra / batches (records without them hash as before).

Answer primitives

  • Answer primitives — every answer is a value and a confidence, declared by types, from plain rules, learned parts and
    model decisions alike (solvi.primitives; guide: "Answer primitives"; examples/16_primitives.py):
    • "Not stated": solvi.Unknown (type NotStated; Maybe[T] = T | NotStated; Answer.maybe(t)) is a real answer —
      the text does not state it — with a confidence, distinct from "no" and from an abstention (None). result.not_stated,
      res.not_stated, res.overall["not_stated"]; constraints see Unknown and joint decoding can choose it; it
      round-trips through JSON ("not_stated": true, probability key "<not stated>").
    • Evidence: Claim(value, evidence=[Quote | str], confidence=, source=) from any part, Decision(..., evidence=) from
      a model; strings are located in the text, every quote must be literally in its given text at its offsets — else the
      output is rejected (safeguard "grounding rejected": the fact is missing, the next producer runs, else the answer
      abstains). Recorded in record.extra["evidence"] (hashed, replayed, tampering caught), result.evidence, shown in the
      audit and counted in the support (quoted / quoted_by_model). Question(require_evidence=True): an answer without a
      quote abstains — the new safeguard evidence missing (guard="evidence_missing", system.stats["evidence_missing"],
      listed by safeguard_report() once it fires).
    • Span: Span[T] / Answer.span(source=, type=) — an exact substring of a given text (a Quote, or a text that is
      located), always grounded, coerced to T with pydantic (a failure: "type rejected"); result.span.
    • *...
Read more