Highlights
- Adds aggregate-only results for a 35-case, 427-call Anthropic tool-grounded expansion.
- Adds a 75-call standalone calibration of the model-based semantic checker.
- Adds a 175-call integrated controller pilot with trigger-level diagnosis.
- Extends
savc verifyand the SHA-256 manifest from four to seven public result files. - Publishes the evaluation fixtures and matching aggregates on Hugging Face.
Main interpretation
The integrated pilot completed all planned calls and preserved all 25 controlled handoffs. The deterministic source-ID check caused all ten observed corrections. The semantic checker allowed all 50 messages it reviewed and caused no correction, so this pilot did not demonstrate an added semantic-intervention benefit. Every condition was already correct on all five action cases, leaving no room to show an action improvement.
These are single-provider development results. They do not establish cross-provider transfer, confirmation-holdout performance, autonomous scientific discovery, or improved drug-development outcomes.
Reproducibility
- 37 tests pass with 81.43% coverage.
- CI passes on Python 3.10, 3.12, 3.13, and 3.14.
- The public-tree and secret scans pass.
- The wheel passes clean-environment
savc demoandsavc verifyruns.
The release contains synthetic fixtures and aggregate outputs only. Raw hosted-model responses, private labels and identifiers, credentials, local paths, and internal execution records are excluded.