Paradigm converts validated deliberative experience into trusted local procedural reflexes, and falls back to the model outside certified regions. v0.1.0 established the Core result. v0.2.0 makes it usable from an application and records what happens when it is plugged into a real agent driving a real model.
No Core benchmark, threshold, or gate was changed in this release.
What is new
paradigm.integration: the public surface. A structuredParadigmState, the actions available now, and aVerifiedOutcomeafter execution.decidereturns a typedReflexDecisionorDeliberateDecision;observeandclose_episodefeed only verified successes into the same acquisition, certification, retention-probe and promotion path measured in P2.1 to P2.4.- Two enforced invariants: an unverified outcome is never positive evidence, and a
DeliberateDecisioncarries the reflex's proposed action for audit but Paradigm never executes it. Capability is not authorization. paradigm serve: the engine as JSON over localhost, with a persisted state file, a per-family trust manifest, per-decision telemetry, and a decision log.- Family-scoped certification (
--certification family_scoped): one compiled artifact, certified family by family against the unchanged criteria. Families that pass become active with their own threshold; the rest stay deliberative. paradigm.integrationplus the LaRuche adapter contract were added to Paradigm. Level 2 scope: a reflex may replay a whitelisted tool with the exact argument template it was validated with, and everything that writes, pushes, delegates, or was never validated stays deliberative. A companion Rust bridge was validated in the LaRuche repository on branchparadigm-integration; it is not part of this Paradigm release tag.examples/integration_minimal.py,docs/INTEGRATION.md, and the full records underresults/integration_laruche/.
Recorded real-provider runs
LaRuche running real missions against deepseek-v4-flash, one bug per mission in a reset workspace, each mission verified independently by pytest rather than by the model's claim of completion.
Runs 9A and 9B, whole-candidate certification, 32 missions: 32 of 32 verified, 0 unsafe actions, 0 false fast paths, and 0 reflexes promoted. The buffer showed why: two families are perfectly consistent (start led to python -m pytest -q 16 times out of 16), but the single tree judged as a whole fails calibration because of the inconsistent post-failure read families, so the consistent ones were never activated.
Run 11, family-scoped activation, fresh state, same model, workspace, guard and thresholds, 24 missions:
missions verified by pytest 24 / 24
first promotion mission 16, version 1
active families laruche:start:none:none, laruche:file_edit:success:other
reflex decisions 15 (all replaying the canonical python -m pytest -q)
model calls avoided 15
false fast paths 0
unsafe actions executed 0
8 of those 15 reflex outcomes were UNKNOWN because of LaRuche shell-output loss, not because the reflex action was observed to fail.
| Segment | Model calls / mission | Tokens / mission | Wall time / mission |
|---|---|---|---|
| Missions 1–16, no reflex | 5.9 | 56.4k | 10.9 s |
| Missions 17–23, v1 active | 4.3 | 37.4k | 7.2 s |
The 17–23 averages are reported as the stable post-activation window. Mission 24 is retained as a recorded outlier and is not used to claim an overall post-activation average improvement. It cost 18 model calls and 74.5 s of self-verification by the model after the same start-of-mission replay, and was still verified; with it the segment average is 6.0 calls, 54.6k tokens, 15.6 s.
Limits of that evidence
- One live run, one model, one workspace, three bug variants. This is an integration record, not a benchmark.
- The
startfamily activated on 5 validated traces. The pre-registered criteria have no per-family minimum count and none was added for the run. - Once a family is served by the reflex it stops producing deliberative held-out evidence, so under the rule as implemented it reports
insufficientat every later compile point and blocks replacement artifacts. The frozen retention probes are the evidence meant for that case; re-certifying an active family on its probes alone is the next refinement and is not in this version. - 8 of the 15 reflex outcomes are UNKNOWN, not SUCCESS, because LaRuche's shell tool drops the output of a failing command. The model's own identical call gets the same.
- Still not demonstrated: argument synthesis, generative content, trust extension into capable-but-unevidenced regions, persistent registry integration, arbitrary domains.
Containment
An early real-provider attempt saw the model install pytest into the user site and create /opt/homebrew/bin/python. Both were reverted, and every recorded run since uses a blocking pre_tool hook that refuses installs, links, copies, moves, deletions, redirections, and any path outside the workspace. The guard exists because of that incident.
Reproduce
pip install -e .
paradigm benchmark p24 --from-cache # recorded Core result, no model endpoint needed
pytest -q # 65 tests, fixture transports only
python examples/integration_minimal.pyRead
- Integration contract and every recorded run:
docs/INTEGRATION.md - Run 11 in full:
results/integration_laruche/run11_family_scoped.md - Offline family audit on the frozen 32-mission record:
results/integration_laruche/family_audit.md - Core result and project state:
results/core_p24/SUMMARY.md,PROJECT_STATE.md,docs/LIMITATIONS.md