更新内容 / What's changed
-
明确低、中、高影响任务的执行档位,以及用户已授权范围内直接执行的规则。
-
增加先列变更路径、排除凭据等受保护路径后再读取安全 diff 的要求。
-
增加 minimal、standard、strict 三种文档治理级别;未配置的既有项目继续采用 standard。
-
将验证报告绑定到 v0.4.0 快照,并补充自然语言评测归因与任务契约限制。
-
Clarified low-, medium-, and high-impact workflows and direct execution within an explicitly authorized scope.
-
Added path-first Git review that excludes protected credential paths before reading safe diffs.
-
Added minimal, standard, and strict documentation-governance levels; existing projects without a setting remain on standard.
-
Bound the verification report to the v0.4.0 snapshot and documented benchmark attribution and task-contract limitations.
2026-09-23 评测说明 / Benchmark note
经用户确认,完成 gpt-6-luna 单模型自然任务评测:7 类任务各 10 对,共 140 次运行。124/140 项任务结果通过,其中 baseline 65/70、with-r-doc 59/70;forbidden-read 为 0/140。70 次 with-r-doc 运行的 Skill visibility 均为 unknown,因此 verified activation 为 0/70、有效配对比较为 0。该结果仅作描述性任务证据,不能归因 r-doc 的独立净收益。
With user confirmation, a single-model gpt-6-luna naturalistic evaluation completed: 7 task classes × 10 pairs, 140 runs. Task outcomes passed 124/140 (baseline 65/70; with-r-doc 59/70); forbidden reads were 0/140. Skill visibility was unknown in all 70 with-r-doc runs, leaving 0/70 verified activations and 0 valid paired comparisons. Treat the results as descriptive task evidence; they do not establish r-doc's independent net benefit.