Releases: riesaexe/r-doc
Release list
r-doc v1.0.0
更新内容 / What's changed
-
明确低、中、高影响任务的执行档位,以及用户已授权范围内直接执行的规则。
-
增加先列变更路径、排除凭据等受保护路径后再读取安全 diff 的要求。
-
增加 minimal、standard、strict 三种文档治理级别;未配置的既有项目继续采用 standard。
-
将验证报告绑定到 v0.4.0 快照,并补充自然语言评测归因与任务契约限制。
-
Clarified low-, medium-, and high-impact workflows and direct execution within an explicitly authorized scope.
-
Added path-first Git review that excludes protected credential paths before reading safe diffs.
-
Added minimal, standard, and strict documentation-governance levels; existing projects without a setting remain on standard.
-
Bound the verification report to the v0.4.0 snapshot and documented benchmark attribution and task-contract limitations.
2026-09-23 评测说明 / Benchmark note
经用户确认,完成 gpt-6-luna 单模型自然任务评测:7 类任务各 10 对,共 140 次运行。124/140 项任务结果通过,其中 baseline 65/70、with-r-doc 59/70;forbidden-read 为 0/140。70 次 with-r-doc 运行的 Skill visibility 均为 unknown,因此 verified activation 为 0/70、有效配对比较为 0。该结果仅作描述性任务证据,不能归因 r-doc 的独立净收益。
With user confirmation, a single-model gpt-6-luna naturalistic evaluation completed: 7 task classes × 10 pairs, 140 runs. Task outcomes passed 124/140 (baseline 65/70; with-r-doc 59/70); forbidden reads were 0/140. Skill visibility was unknown in all 70 with-r-doc runs, leaving 0/70 verified activations and 0 valid paired comparisons. Treat the results as descriptive task evidence; they do not establish r-doc's independent net benefit.
r-doc v0.4.1
r-doc v0.4.1
- 收敛 naturalistic grader 契约,补充 v2 复杂任务规格、批次编排、公开证据校验和可观察耗时/token 指标。
- 完成 gpt-5.6-luna 新 runner 复采:14-run smoke 先行通过;完整 A/B 共 140 次运行,133 次 task outcome 通过、7 次失败,0 个 runner 异常,140/140 激活证据验证通过。
- 公开证据校验通过;结果受两组共享 safe-read preflight 影响,仅作 descriptive-only 证据,不宣称 r-doc 独立增量收益。
- 强化索引目录文档治理,并修正版本化任务根目录下的 manifest provenance。
- 为 byte-hashed benchmark 文本固定 UTF-8 LF checkout,避免 Windows 原始字节重放差异。
Package SHA-256: ff1f7352fba2b21d821769fd8284809bf501e32095bdd19a7a5e649aabba45db
English
- Tightened naturalistic grader contracts, added v2 complex-task specifications, batch orchestration, public-evidence verification, and observable duration/token metrics.
- Completed the new
gpt-5.6-lunarunner recapture: a 14-run smoke gate passed first; the full A/B batch contains 140 runs, 133 task-outcome passes, 7 failures, zero runner exceptions, and 140/140 verified activation records. - Public-evidence verification passes. Because both conditions share the safe-read preflight, this remains descriptive-only evidence and is not an independent r-doc-effect claim.
- Strengthened indexed-document governance and corrected manifest provenance for versioned task roots.
- Fixed byte-hashed benchmark text to LF at checkout to avoid Windows raw-byte replay differences.
Package SHA-256: ff1f7352fba2b21d821769fd8284809bf501e32095bdd19a7a5e649aabba45db
r-doc v0.4.0
Highlights
- Adds bounded, deny-by-default documentation reads with protected secret paths and explicit search scopes.
- Adds decision-note lifecycle, validation, supersession, and preview-first archive governance.
- Publishes the naturalistic evidence chain with per-run activation signals, schema-2 hashes, and sanitized public evidence.
- Fixes the Event v2 benchmark static-assertion false negative; the immutable original v8 result remains available alongside the corrected 80/80 derived regrade.
- Requires local Git Credential Manager plus GitHub REST API for release creation and upload.
Verification
- 128 unit tests passed.
- Skill package validation passed.
- Strict documentation audit passed with zero errors and warnings.
- Agent evidence evaluation passed at 100%.
- Signed commit and tag: 12775bc / v0.4.0.
Projects without a configured decision-note root remain compatible. Naturalistic results are runner/fixture evidence, not a direct Skill-effect claim.
r-doc v0.3.0
Highlights
- Added optional decision-note governance with templates, CLI support, and audit integration.
- Added preview-first archive planning and safe repair workflow.
- Added benchmark governance and portable aggregate paths.
Verification
- 116 unit tests passed.
- Skill package validator, strict audit, repair preview, official validator, and local npx discovery passed.
- User-approved gpt-5.6-luna naturalistic batch: 80 measurement-valid runs, 40 pairs, 10 pairs per task across 4 tasks; 76 forbidden-read failures and 5% task-success/context-safety in both conditions. This is not a positive effectiveness claim.
Compatibility and limitations
- Existing projects without a configured decision-note root remain compatible.
- Public repository includes the Luna aggregate summary only; raw CLI traces/logs remain local because they contain machine-specific paths.
- Package: r-doc-v0.3.0.zip.
r-doc v0.2.17
English
r-doc v0.2.17 closes the independent naturalistic capture boundary.
- The runner builds isolated fixtures, invokes the agent with only the user task, independently snapshots the final workspace, normalizes the raw CLI trace, hashes capture artifacts, and invokes the independent grader.
- The grader now owns executable pytest, callable-behavior, and documentation outcome checks.
- Windows path normalization rejects the .\env forbidden-read bypass, and the conformance aggregator rejects runs mislabeled as naturalistic.
- Four naturalistic task specifications and a separate multi-run aggregator are included.
- The first real evidence set contains eight Codex gpt-5.5 captures across four matched A/B pairs. It is explicitly partial: all runs fail context safety on forbidden .env/secrets.md reads, so this release makes no positive naturalistic effectiveness claim.
- Historical release notes from v0.1.0 through v0.2.16 now have complete English and Chinese sections.
Verification
- validate_skill.py: PASS
- strict documentation audit: PASS, 0 errors and 0 warnings
- pytest: 96 passed
- unittest: 96 passed
- strict Agent evaluation: 100 percent
- npx skills add . --list: r-doc discovered
- release commit: 9931fc1, GPG-signed locally
中文
r-doc v0.2.17 闭合了独立 naturalistic capture 边界。
- runner 构建隔离 fixture,只向 Agent 发送用户任务,独立读取最终 workspace、规范化原始 CLI trace、哈希捕获产物,并调用独立 grader。
- grader 现在自己执行 pytest、callable 行为和文档 outcome 检查。
- Windows 路径规范化拒绝 .\env 禁止读取侧门,conformance 聚合器拒绝误标为 naturalistic 的运行记录。
- 增加四个 naturalistic 任务规范和独立多运行聚合器。
- 首批真实证据包含四个匹配 A/B 对、共 8 次 Codex gpt-5.5 capture。证据明确为 partial:所有运行都因读取禁止的 .env/secrets.md 而 context safety 失败,因此本版本不宣称正向 naturalistic effectiveness result。
- v0.1.0 至 v0.2.16 的历史 Release notes 现在全部包含完整中英文段落。
验证
- validate_skill.py:通过
- 严格文档审计:通过,0 error、0 warning
- pytest:96 passed
- unittest:96 passed
- 严格 Agent 评测:100%
- npx skills add . --list:发现 r-doc
- 发布提交:9931fc1,本地 GPG 签名
r-doc v0.2.16
English
r-doc v0.2.16 establishes the bilingual public-document baseline and formalizes the two-layer benchmark model.
Highlights
- Added Chinese summaries to every historical section in
CHANGELOG.md. - Added bilingual public release, development, architecture, and benchmark documentation.
- Reclassified fixed-prompt runs as Conformance Benchmark / skill-layer ablation.
- Added machine-readable benchmark provenance metadata and the independent Naturalistic Effectiveness protocol/grader.
- Kept the benchmark interpretation explicit: three matched pairs are trend-ready only; no naturalistic effectiveness result is claimed.
- Enabled and retained protected
maingovernance.
Verification
validate_skill.py: PASSaudit_docs.py --strict: PASS, 0 errors / 0 warnings- Unit suite: 85 tests passed
- Agent evidence evaluator: PASS
- Benchmark aggregation: PASS, current summary remains
partial - Signed release tag:
v0.2.16
The historical v0.2.15 commit remains unchanged and unsigned. The v0.2.16 tag points to a locally GPG-signed commit; GitHub's Verified display should be checked on the tag commit.
中文
r-doc v0.2.16 建立公开文档双语基线,并正式明确两层 benchmark 模型。
主要变化
- 为
CHANGELOG.md的所有历史版本增加中文摘要。 - 为公开发布、开发、架构和 benchmark 文档增加中英双语内容。
- 将固定 prompt 运行重新命名为 Conformance Benchmark / skill-layer ablation。
- 增加机器可读的 benchmark provenance 元数据,以及独立的 Naturalistic Effectiveness 协议和 grader。
- 明确 benchmark 边界:当前 3 组匹配运行仅达到趋势就绪,不宣称 naturalistic effectiveness 结果。
- 启用并保留
main分支保护。
验证
validate_skill.py:PASSaudit_docs.py --strict:PASS,0 errors / 0 warnings- 单元测试:85 个通过
- Agent evidence evaluator:PASS
- benchmark 聚合:PASS,当前 summary 仍为
partial - 签名发布 tag:
v0.2.16
历史 v0.2.15 commit 保持原样且仍为 unsigned。v0.2.16 tag 指向本机 GPG 签名 commit;仍应在 GitHub 页面确认该 tag commit 显示 Verified。
r-doc v0.2.15
English
Highlights
- Upgraded evidence to schema v3 and trace validation to schema v2, with trace-derived prompt, activation, selected-skill, report, response, diff, and review evidence.
- Added condition-aware paired benchmark deltas, three readiness levels, and descriptive 95% Student-t intervals.
- Added six valid official benchmark runs and three paired local Codex gpt-5.5 comparisons. The data is trend-ready but not statistical-ready.
- Split required, allowed, and forbidden reads, and made forbidden reads fail strict evaluation and benchmark aggregation gates.
- Aligned the public changelog with the English README baseline and added discoverability topics to the repository.
Validation
- Agent evidence strict evaluation: 100% (56/56).
- Skill package validation: PASS.
- Strict documentation audit: PASS, 0 errors and 0 warnings.
- Regression suite: 80 tests passed.
- Local npx skills add . --list: r-doc discovered successfully.
Benchmark note
The checked-in real-agent data supports observing behavior trends. With three pairs, it does not support a claim that r-doc reliably improves task success; more paired runs and model coverage are still required.
中文
- 将 evidence schema 升级到 v3,并要求 trace 独立记录 prompt、activation、skill、报告、最终响应、diff 和 review,再由聚合器交叉校验。
- 将 trace schema 升级到 v2,按事件类型拒绝未声明字段;forbidden read 在普通评估中降低 context economy,在 strict 和聚合门禁中失败。
- 增加配对 delta 的 trend/statistical/strong-evidence 样本门槛、95% Student-t 区间,以及真实 Codex 捕获记录和运行说明。
- 澄清 condition 与 Skill 选择的边界,保留失败 capture 供审计但排除出聚合,并强制 Windows CLI 使用 UTF-8 保存 CJK/emoji trace。
r-doc v0.2.14
English
[0.2.14] - 2026-09-15
- 将真实 Agent benchmark 的 trace 升级为带运行元数据绑定、场景生命周期和 action 事件的结构化 JSONL,并从 trace 交叉验证 paths、读取、命令和写入证据。
- 为 benchmark 聚合增加按
agent + model + run_id配对的with-r-doc/baseline-no-r-doc差值、均值、中位数、标准差和统计就绪度;profile 汇总不再混合 condition。 - 将 context economy 的读取策略拆分为 required、allowed 和 forbidden 集合,分别报告 unnecessary、forbidden 和缺失 required reads。
- 将审计性能基线默认迭代次数提高到 10,并同时记录线性插值 p95、最大值和低样本提示。
中文
- 将真实 Agent benchmark trace 升级为带运行元数据绑定、场景生命周期和 action 事件的结构化 JSONL,并从 trace 交叉验证路径、读取、命令和写入证据。
- 按
agent + model + run_id输出with-r-doc与baseline-no-r-doc配对 delta,并提供均值、中位数、标准差和统计就绪度;profile 不再混合 condition。 - 将 context-economy 读取策略拆为 required、allowed、forbidden,分别报告 unnecessary、forbidden 和 missing-required reads。
- 审计性能基线默认提升到 10 次迭代,并记录插值 p95、最大值和低样本提示。
r-doc v0.2.13
English
Highlights
- Added paired Chinese public references for the safe-repair guide, practical examples, and common pitfalls.
- Routed
README.mdto English references andREADME.zh-CN.mdto matching.zh-CN.mdreferences. - Added a regression guard for README language routing and cross-links from localized references to their English canonical files.
- Updated the Chinese Skill reference pointer and release/migration documentation.
Verification
- 68 regression tests passed.
- Skill package validation and the official Skill validator passed.
- Repair preview reported no pending safe repairs.
- Strict audit passed with 0 errors and 0 warnings; only the three reviewed allowlisted-example informational findings remained.
- Complete Agent evidence example passed at 100.0%.
中文
- 为 GitHub 公开入口增加安全修复、实际案例和常见避坑指南的中英文配对文档,并让中文 README 指向中文指南。
- 为 canonical English references 增加中文回链,并让中文 Skill 参考入口使用一致的语言路由。
- 增加公开文档语言路由回归测试,防止 README 链接再次漂移。
r-doc v0.2.12
English
Highlights
- Added the real Agent benchmark directory contract and aggregation pipeline without claiming synthetic evidence as real results.
- Bound
machine_rulesidentifiers bidirectionally to the evaluator code registry. - Added local audit performance baselines for 100, 1000, and 5000 Markdown documents.
Verification
- 66 regression tests passed.
- Skill package validation, repair preview, strict audit, and the official Skill validator passed.
- Benchmark aggregation remains explicitly
pendinguntil captured real-agent traces and a no-r-doc baseline are available.
中文
- 增加真实 Agent benchmark 目录契约、运行 evidence 聚合器和 baseline 说明,不把示例 evidence 当作真实结果。
- 增加
machine_rules标识、代码注册表和cases.json的双向一致性检查,防止规则名称漂移。 - 增加针对 100、1,000、5,000 个 Markdown 文档的审计性能基线工具及首份本机测量记录。