Skip to content

Paradigm v0.2.0 - Integration surface and first reflex promoted in a real agent loop

Latest

Choose a tag to compare

@infinition infinition released this 16 Sep 22:08
· 92 commits to main since this release

Paradigm converts validated deliberative experience into trusted local procedural reflexes, and falls back to the model outside certified regions. v0.1.0 established the Core result. v0.2.0 makes it usable from an application and records what happens when it is plugged into a real agent driving a real model.

No Core benchmark, threshold, or gate was changed in this release.

What is new

  • paradigm.integration: the public surface. A structured ParadigmState, the actions available now, and a VerifiedOutcome after execution. decide returns a typed ReflexDecision or DeliberateDecision; observe and close_episode feed only verified successes into the same acquisition, certification, retention-probe and promotion path measured in P2.1 to P2.4.
  • Two enforced invariants: an unverified outcome is never positive evidence, and a DeliberateDecision carries the reflex's proposed action for audit but Paradigm never executes it. Capability is not authorization.
  • paradigm serve: the engine as JSON over localhost, with a persisted state file, a per-family trust manifest, per-decision telemetry, and a decision log.
  • Family-scoped certification (--certification family_scoped): one compiled artifact, certified family by family against the unchanged criteria. Families that pass become active with their own threshold; the rest stay deliberative.
  • paradigm.integration plus the LaRuche adapter contract were added to Paradigm. Level 2 scope: a reflex may replay a whitelisted tool with the exact argument template it was validated with, and everything that writes, pushes, delegates, or was never validated stays deliberative. A companion Rust bridge was validated in the LaRuche repository on branch paradigm-integration; it is not part of this Paradigm release tag.
  • examples/integration_minimal.py, docs/INTEGRATION.md, and the full records under results/integration_laruche/.

Recorded real-provider runs

LaRuche running real missions against deepseek-v4-flash, one bug per mission in a reset workspace, each mission verified independently by pytest rather than by the model's claim of completion.

Runs 9A and 9B, whole-candidate certification, 32 missions: 32 of 32 verified, 0 unsafe actions, 0 false fast paths, and 0 reflexes promoted. The buffer showed why: two families are perfectly consistent (start led to python -m pytest -q 16 times out of 16), but the single tree judged as a whole fails calibration because of the inconsistent post-failure read families, so the consistent ones were never activated.

Run 11, family-scoped activation, fresh state, same model, workspace, guard and thresholds, 24 missions:

missions verified by pytest   24 / 24
first promotion               mission 16, version 1
active families               laruche:start:none:none, laruche:file_edit:success:other
reflex decisions              15 (all replaying the canonical python -m pytest -q)
model calls avoided           15
false fast paths              0
unsafe actions executed       0

8 of those 15 reflex outcomes were UNKNOWN because of LaRuche shell-output loss, not because the reflex action was observed to fail.

Segment Model calls / mission Tokens / mission Wall time / mission
Missions 1–16, no reflex 5.9 56.4k 10.9 s
Missions 17–23, v1 active 4.3 37.4k 7.2 s

The 17–23 averages are reported as the stable post-activation window. Mission 24 is retained as a recorded outlier and is not used to claim an overall post-activation average improvement. It cost 18 model calls and 74.5 s of self-verification by the model after the same start-of-mission replay, and was still verified; with it the segment average is 6.0 calls, 54.6k tokens, 15.6 s.

Limits of that evidence

  • One live run, one model, one workspace, three bug variants. This is an integration record, not a benchmark.
  • The start family activated on 5 validated traces. The pre-registered criteria have no per-family minimum count and none was added for the run.
  • Once a family is served by the reflex it stops producing deliberative held-out evidence, so under the rule as implemented it reports insufficient at every later compile point and blocks replacement artifacts. The frozen retention probes are the evidence meant for that case; re-certifying an active family on its probes alone is the next refinement and is not in this version.
  • 8 of the 15 reflex outcomes are UNKNOWN, not SUCCESS, because LaRuche's shell tool drops the output of a failing command. The model's own identical call gets the same.
  • Still not demonstrated: argument synthesis, generative content, trust extension into capable-but-unevidenced regions, persistent registry integration, arbitrary domains.

Containment

An early real-provider attempt saw the model install pytest into the user site and create /opt/homebrew/bin/python. Both were reverted, and every recorded run since uses a blocking pre_tool hook that refuses installs, links, copies, moves, deletions, redirections, and any path outside the workspace. The guard exists because of that incident.

Reproduce

pip install -e .
paradigm benchmark p24 --from-cache   # recorded Core result, no model endpoint needed
pytest -q                             # 65 tests, fixture transports only
python examples/integration_minimal.py

Read

  • Integration contract and every recorded run: docs/INTEGRATION.md
  • Run 11 in full: results/integration_laruche/run11_family_scoped.md
  • Offline family audit on the frozen 32-mission record: results/integration_laruche/family_audit.md
  • Core result and project state: results/core_p24/SUMMARY.md, PROJECT_STATE.md, docs/LIMITATIONS.md