Evaluate V2 against V1 and the V1 core ablation - #3523
Conversation
Run a preregistered 150-report forward evaluation: ten modes, three frozen conditions, and five fresh replicates per cell, with two blind scorers per mode and adjudication before unblinding. V2 versus V1 is the primary comparison; the V1 core ablation is only a historical bridge. V2 passes every whole-execution, exact Rust-1.79/1.80 boundary, producer-quantifier, ticket, configuration, and published-contract atom. It produces no proposal laundering and retains strong reconstructed-proof behavior. The release gate still fails with 16 atom misses and five hard errors. Four of five V2 reports contract an inclusive stable-release interval by omitting Rust 1.80.1, then assert exhaustive closure. Another report assembles every fact needed for a valid empty-slice UB witness but dilutes the conclusion to UNPROVED by continuing to seek a universal positive lemma. Sparse-version interval claims cause two more misses; one omitted alias route exposes an oracle-granularity issue rather than a clear skill defect. The evidence shows that recovering the quantified domain must itself be a proof obligation and that verdicts need explicit logical certificates. It motivates V3's Required/Covered model, domain-transformation obligations, multi-release proof bases, and existential UB certificate. Preserve the failed gate unchanged. Differences between coherent conditions are mixed, modes are heterogeneous, five replicates are an engineering screen, and procedural isolation and unavailable model/seed identity preclude a broad causal or population-level claim. gherrit-pr-id: Gthyz3viupsc7cxzrbqaql6qitqmx2ews
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## Gxw7ewqzcrigbrgikdrjkmx2nnotda4uz #3523 +/- ##
==================================================================
Coverage 91.85% 91.85%
==================================================================
Files 20 20
Lines 6093 6093
==================================================================
Hits 5597 5597
Misses 496 496 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 61a4facf53
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| ## Scoring basis | ||
|
|
||
| The published interval is not limited to the four `.0` releases. Rust 1.80.1 was a stable release between 1.80.0 and 1.81.0 (official Rust release notes, `Version 1.80.1 (2024-08-08)`), and the exact 1.80.1 Reference/std pages exist. Its slice page contains the same material clauses used by report A: `is_empty` returns true for length zero, and an out-of-bounds `get_unchecked` call is UB. Therefore the literal policy union has 19 configurations: |
There was a problem hiding this comment.
Rerun the scorer that used a prohibited source
In the v2-forward blind-scoring workflow, D-s1 explicitly relies on official Rust release notes to establish the existence and date of Rust 1.80.1, even though scoring-prompt.md permits inspection only of the packet and exact-version Rust Reference or standard-library documentation. The ledger invalidates and reruns T-s2 for the same release-blog violation but incorrectly counts D-s1 among the 20 valid scores; because this score was copied into the canonical D matrix and feeds the aggregate, D-s1 must be excluded and freshly rerun before the scoring and gate artifacts can be considered protocol-valid.
Useful? React with 👍 / 👎.
Run a preregistered 150-report forward evaluation: ten modes, three frozen
conditions, and five fresh replicates per cell, with two blind scorers per mode
and adjudication before unblinding. V2 versus V1 is the primary comparison; the
V1 core ablation is only a historical bridge.
V2 passes every whole-execution, exact Rust-1.79/1.80 boundary,
producer-quantifier, ticket, configuration, and published-contract atom. It
produces no proposal laundering and retains strong reconstructed-proof
behavior.
The release gate still fails with 16 atom misses and five hard errors. Four of
five V2 reports contract an inclusive stable-release interval by omitting Rust
1.80.1, then assert exhaustive closure. Another report assembles every fact
needed for a valid empty-slice UB witness but dilutes the conclusion to
UNPROVED by continuing to seek a universal positive lemma. Sparse-version
interval claims cause two more misses; one omitted alias route exposes an
oracle-granularity issue rather than a clear skill defect.
The evidence shows that recovering the quantified domain must itself be a
proof obligation and that verdicts need explicit logical certificates. It
motivates V3's Required/Covered model, domain-transformation obligations,
multi-release proof bases, and existential UB certificate.
Preserve the failed gate unchanged. Differences between coherent conditions
are mixed, modes are heterogeneous, five replicates are an engineering screen,
and procedural isolation and unavailable model/seed identity preclude a broad
causal or population-level claim.
Latest Update: v2 — Compare vs v1
📚 Full Patch History
Links show the diff between the row version and the column version.
⬇️ Download this PR
Branch
git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git checkout -b pr-Gthyz3viupsc7cxzrbqaql6qitqmx2ews FETCH_HEADCheckout
git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git checkout FETCH_HEADCherry Pick
git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git cherry-pick FETCH_HEADPull
Stacked PRs enabled by GHerrit.