Releases: infinition/Paradigm
Release list
Paradigm v0.2.0 - Integration surface and first reflex promoted in a real agent loop
Paradigm converts validated deliberative experience into trusted local procedural reflexes, and falls back to the model outside certified regions. v0.1.0 established the Core result. v0.2.0 makes it usable from an application and records what happens when it is plugged into a real agent driving a real model.
No Core benchmark, threshold, or gate was changed in this release.
What is new
paradigm.integration: the public surface. A structuredParadigmState, the actions available now, and aVerifiedOutcomeafter execution.decidereturns a typedReflexDecisionorDeliberateDecision;observeandclose_episodefeed only verified successes into the same acquisition, certification, retention-probe and promotion path measured in P2.1 to P2.4.- Two enforced invariants: an unverified outcome is never positive evidence, and a
DeliberateDecisioncarries the reflex's proposed action for audit but Paradigm never executes it. Capability is not authorization. paradigm serve: the engine as JSON over localhost, with a persisted state file, a per-family trust manifest, per-decision telemetry, and a decision log.- Family-scoped certification (
--certification family_scoped): one compiled artifact, certified family by family against the unchanged criteria. Families that pass become active with their own threshold; the rest stay deliberative. paradigm.integrationplus the LaRuche adapter contract were added to Paradigm. Level 2 scope: a reflex may replay a whitelisted tool with the exact argument template it was validated with, and everything that writes, pushes, delegates, or was never validated stays deliberative. A companion Rust bridge was validated in the LaRuche repository on branchparadigm-integration; it is not part of this Paradigm release tag.examples/integration_minimal.py,docs/INTEGRATION.md, and the full records underresults/integration_laruche/.
Recorded real-provider runs
LaRuche running real missions against deepseek-v4-flash, one bug per mission in a reset workspace, each mission verified independently by pytest rather than by the model's claim of completion.
Runs 9A and 9B, whole-candidate certification, 32 missions: 32 of 32 verified, 0 unsafe actions, 0 false fast paths, and 0 reflexes promoted. The buffer showed why: two families are perfectly consistent (start led to python -m pytest -q 16 times out of 16), but the single tree judged as a whole fails calibration because of the inconsistent post-failure read families, so the consistent ones were never activated.
Run 11, family-scoped activation, fresh state, same model, workspace, guard and thresholds, 24 missions:
missions verified by pytest 24 / 24
first promotion mission 16, version 1
active families laruche:start:none:none, laruche:file_edit:success:other
reflex decisions 15 (all replaying the canonical python -m pytest -q)
model calls avoided 15
false fast paths 0
unsafe actions executed 0
8 of those 15 reflex outcomes were UNKNOWN because of LaRuche shell-output loss, not because the reflex action was observed to fail.
| Segment | Model calls / mission | Tokens / mission | Wall time / mission |
|---|---|---|---|
| Missions 1–16, no reflex | 5.9 | 56.4k | 10.9 s |
| Missions 17–23, v1 active | 4.3 | 37.4k | 7.2 s |
The 17–23 averages are reported as the stable post-activation window. Mission 24 is retained as a recorded outlier and is not used to claim an overall post-activation average improvement. It cost 18 model calls and 74.5 s of self-verification by the model after the same start-of-mission replay, and was still verified; with it the segment average is 6.0 calls, 54.6k tokens, 15.6 s.
Limits of that evidence
- One live run, one model, one workspace, three bug variants. This is an integration record, not a benchmark.
- The
startfamily activated on 5 validated traces. The pre-registered criteria have no per-family minimum count and none was added for the run. - Once a family is served by the reflex it stops producing deliberative held-out evidence, so under the rule as implemented it reports
insufficientat every later compile point and blocks replacement artifacts. The frozen retention probes are the evidence meant for that case; re-certifying an active family on its probes alone is the next refinement and is not in this version. - 8 of the 15 reflex outcomes are UNKNOWN, not SUCCESS, because LaRuche's shell tool drops the output of a failing command. The model's own identical call gets the same.
- Still not demonstrated: argument synthesis, generative content, trust extension into capable-but-unevidenced regions, persistent registry integration, arbitrary domains.
Containment
An early real-provider attempt saw the model install pytest into the user site and create /opt/homebrew/bin/python. Both were reverted, and every recorded run since uses a blocking pre_tool hook that refuses installs, links, copies, moves, deletions, redirections, and any path outside the workspace. The guard exists because of that incident.
Reproduce
pip install -e .
paradigm benchmark p24 --from-cache # recorded Core result, no model endpoint needed
pytest -q # 65 tests, fixture transports only
python examples/integration_minimal.pyRead
- Integration contract and every recorded run:
docs/INTEGRATION.md - Run 11 in full:
results/integration_laruche/run11_family_scoped.md - Offline family audit on the frozen 32-mission record:
results/integration_laruche/family_audit.md - Core result and project state:
results/core_p24/SUMMARY.md,PROJECT_STATE.md,docs/LIMITATIONS.md
Paradigm v0.1.0 - First Type B Procedural Memory Baseline
Paradigm converts validated deliberative experience into trusted local procedural reflexes, while falling back to the deliberative model outside certified regions. An agent keeps using its language model for anything new; when a decision has been made successfully often enough, Paradigm compiles it into a small local reflex, certifies it, and only then lets the reflex act. Unknown or uncertified situations stay with the model.
This release is the first versioned baseline with a Type B result: a procedure that did not exist in the system was learned from a handful of validated model demonstrations, compiled into a reflex, promoted during online use, and executed on fresh tasks without the model.
The result in three lines
Frozen evaluation on 8 held-out dependency tasks:
LLM-only 6/8
Hybrid 6/8
Reflex-only 8/8
The hybrid system stayed at 6 of 8 because the trust gate correctly rejected the two prompt variants for which no validated evidence existed and fell back to the teacher, which failed on them. The frozen reflex itself generalized successfully to both. It saved 61% of LLM calls (22 against 56) and 57% of tokens on those tasks with zero false fast paths.
Demonstrated in this release
- Replicated Type A certified coverage expansion: a previously unrepresented state region becomes certified online from validated experience, reproduced across 3 seeds, 4 arrival orders, and 2 teacher models with zero seed variance (P2.3R).
- Genuine Type B procedural capability acquisition: capability 50% with zero novel evidence, 100% after 4 to 6 validated demonstrations, Time-to-Capability median 4 episodes (P2.4).
- Frozen reflex generalization on held-out tasks, including two variants the teacher fails on (P2.4).
- Retention-probe protection: immutable per-family probes catch a candidate that damages a mature family and that the recent-buffer rule promotes (P2.3R-bis).
- Online promotion and fallback: the reflex is promoted at episode 50 of a 69-episode stream and solves later episodes with zero LLM calls; uncertified states stay deliberative; 0 invalid reflex actions and 0 false fast paths in every recorded run.
- Measured LLM-call and token savings: 61% fewer LLM calls online and 61% fewer on the frozen held-out set, at task success equal to the model alone.
Not demonstrated
- Universal model independence. Only two Qwen sizes were used; the larger one supplied too few validated demonstrations to learn the Type B procedure.
- Arbitrary-domain procedural learning. One controlled coding family, one online seed.
- Robotics or world-model results.
- Safe trust expansion into regions without validated evidence.
- Type C compositional acquisition.
Open research question
How can Paradigm learn where an acquired reflex can be trusted without lowering its safety gate? The reflex solved 8 of 8 held-out tasks; Paradigm had evidence to trust it on 6.
Reproduce
pip install -e .
paradigm benchmark p24 --from-cache # recorded result, no model endpoint needed
pytest -q # 50 tests, fixture transports onlyA live run needs an OpenAI-compatible or native Ollama endpoint; see benchmarks/README.md.
Read
- One page:
results/core_p24/SUMMARY.md - Full reading:
results/core_p24/INTERPRETATION.md - Replication line:
results/core_p23r/INTERPRETATION.md,results/core_p23r_bis/INTERPRETATION.md,results/core_p23t/INTERPRETATION.md - Project state and limitations:
PROJECT_STATE.md,docs/LIMITATIONS.md
The immutable scientific snapshot is tag p2.4-type-b-baseline (commit 78e0a25); this release adds a Python 3.11 syntax fix, the paradigm CLI, and the one-page summary on top of identical recorded artifacts.