Paradigm v0.1.0 - First Type B Procedural Memory Baseline
Paradigm converts validated deliberative experience into trusted local procedural reflexes, while falling back to the deliberative model outside certified regions. An agent keeps using its language model for anything new; when a decision has been made successfully often enough, Paradigm compiles it into a small local reflex, certifies it, and only then lets the reflex act. Unknown or uncertified situations stay with the model.
This release is the first versioned baseline with a Type B result: a procedure that did not exist in the system was learned from a handful of validated model demonstrations, compiled into a reflex, promoted during online use, and executed on fresh tasks without the model.
The result in three lines
Frozen evaluation on 8 held-out dependency tasks:
LLM-only 6/8
Hybrid 6/8
Reflex-only 8/8
The hybrid system stayed at 6 of 8 because the trust gate correctly rejected the two prompt variants for which no validated evidence existed and fell back to the teacher, which failed on them. The frozen reflex itself generalized successfully to both. It saved 61% of LLM calls (22 against 56) and 57% of tokens on those tasks with zero false fast paths.
Demonstrated in this release
- Replicated Type A certified coverage expansion: a previously unrepresented state region becomes certified online from validated experience, reproduced across 3 seeds, 4 arrival orders, and 2 teacher models with zero seed variance (P2.3R).
- Genuine Type B procedural capability acquisition: capability 50% with zero novel evidence, 100% after 4 to 6 validated demonstrations, Time-to-Capability median 4 episodes (P2.4).
- Frozen reflex generalization on held-out tasks, including two variants the teacher fails on (P2.4).
- Retention-probe protection: immutable per-family probes catch a candidate that damages a mature family and that the recent-buffer rule promotes (P2.3R-bis).
- Online promotion and fallback: the reflex is promoted at episode 50 of a 69-episode stream and solves later episodes with zero LLM calls; uncertified states stay deliberative; 0 invalid reflex actions and 0 false fast paths in every recorded run.
- Measured LLM-call and token savings: 61% fewer LLM calls online and 61% fewer on the frozen held-out set, at task success equal to the model alone.
Not demonstrated
- Universal model independence. Only two Qwen sizes were used; the larger one supplied too few validated demonstrations to learn the Type B procedure.
- Arbitrary-domain procedural learning. One controlled coding family, one online seed.
- Robotics or world-model results.
- Safe trust expansion into regions without validated evidence.
- Type C compositional acquisition.
Open research question
How can Paradigm learn where an acquired reflex can be trusted without lowering its safety gate? The reflex solved 8 of 8 held-out tasks; Paradigm had evidence to trust it on 6.
Reproduce
pip install -e .
paradigm benchmark p24 --from-cache # recorded result, no model endpoint needed
pytest -q # 50 tests, fixture transports onlyA live run needs an OpenAI-compatible or native Ollama endpoint; see benchmarks/README.md.
Read
- One page:
results/core_p24/SUMMARY.md - Full reading:
results/core_p24/INTERPRETATION.md - Replication line:
results/core_p23r/INTERPRETATION.md,results/core_p23r_bis/INTERPRETATION.md,results/core_p23t/INTERPRETATION.md - Project state and limitations:
PROJECT_STATE.md,docs/LIMITATIONS.md
The immutable scientific snapshot is tag p2.4-type-b-baseline (commit 78e0a25); this release adds a Python 3.11 syntax fix, the paradigm CLI, and the one-page summary on top of identical recorded artifacts.