r-doc v0.2.15
English
Highlights
- Upgraded evidence to schema v3 and trace validation to schema v2, with trace-derived prompt, activation, selected-skill, report, response, diff, and review evidence.
- Added condition-aware paired benchmark deltas, three readiness levels, and descriptive 95% Student-t intervals.
- Added six valid official benchmark runs and three paired local Codex gpt-5.5 comparisons. The data is trend-ready but not statistical-ready.
- Split required, allowed, and forbidden reads, and made forbidden reads fail strict evaluation and benchmark aggregation gates.
- Aligned the public changelog with the English README baseline and added discoverability topics to the repository.
Validation
- Agent evidence strict evaluation: 100% (56/56).
- Skill package validation: PASS.
- Strict documentation audit: PASS, 0 errors and 0 warnings.
- Regression suite: 80 tests passed.
- Local npx skills add . --list: r-doc discovered successfully.
Benchmark note
The checked-in real-agent data supports observing behavior trends. With three pairs, it does not support a claim that r-doc reliably improves task success; more paired runs and model coverage are still required.
中文
- 将 evidence schema 升级到 v3,并要求 trace 独立记录 prompt、activation、skill、报告、最终响应、diff 和 review,再由聚合器交叉校验。
- 将 trace schema 升级到 v2,按事件类型拒绝未声明字段;forbidden read 在普通评估中降低 context economy,在 strict 和聚合门禁中失败。
- 增加配对 delta 的 trend/statistical/strong-evidence 样本门槛、95% Student-t 区间,以及真实 Codex 捕获记录和运行说明。
- 澄清 condition 与 Skill 选择的边界,保留失败 capture 供审计但排除出聚合,并强制 Windows CLI 使用 UTF-8 保存 CJK/emoji trace。