Skip to content

Evaluate V2 against V1 and the V1 core ablation - #3523

Open
joshlf wants to merge 1 commit into
Gxw7ewqzcrigbrgikdrjkmx2nnotda4uzfrom
Gthyz3viupsc7cxzrbqaql6qitqmx2ews
Open

Evaluate V2 against V1 and the V1 core ablation#3523
joshlf wants to merge 1 commit into
Gxw7ewqzcrigbrgikdrjkmx2nnotda4uzfrom
Gthyz3viupsc7cxzrbqaql6qitqmx2ews

Conversation

@joshlf

@joshlf joshlf commented Aug 3, 2026

Copy link
Copy Markdown
Member

Run a preregistered 150-report forward evaluation: ten modes, three frozen
conditions, and five fresh replicates per cell, with two blind scorers per mode
and adjudication before unblinding. V2 versus V1 is the primary comparison; the
V1 core ablation is only a historical bridge.

V2 passes every whole-execution, exact Rust-1.79/1.80 boundary,
producer-quantifier, ticket, configuration, and published-contract atom. It
produces no proposal laundering and retains strong reconstructed-proof
behavior.

The release gate still fails with 16 atom misses and five hard errors. Four of
five V2 reports contract an inclusive stable-release interval by omitting Rust
1.80.1, then assert exhaustive closure. Another report assembles every fact
needed for a valid empty-slice UB witness but dilutes the conclusion to
UNPROVED by continuing to seek a universal positive lemma. Sparse-version
interval claims cause two more misses; one omitted alias route exposes an
oracle-granularity issue rather than a clear skill defect.

The evidence shows that recovering the quantified domain must itself be a
proof obligation and that verdicts need explicit logical certificates. It
motivates V3's Required/Covered model, domain-transformation obligations,
multi-release proof bases, and existential UB certificate.

Preserve the failed gate unchanged. Differences between coherent conditions
are mixed, modes are heterogeneous, five replicates are an engineering screen,
and procedural isolation and unavailable model/seed identity preclude a broad
causal or population-level claim.


Latest Update: v2 — Compare vs v1

📚 Full Patch History

Links show the diff between the row version and the column version.

Version v1 Base
v2 vs v1 vs Base
v1 vs Base
⬇️ Download this PR

Branch

git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git checkout -b pr-Gthyz3viupsc7cxzrbqaql6qitqmx2ews FETCH_HEAD

Checkout

git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git checkout FETCH_HEAD

Cherry Pick

git fetch origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews && git cherry-pick FETCH_HEAD

Pull

git pull origin refs/heads/Gthyz3viupsc7cxzrbqaql6qitqmx2ews

Stacked PRs enabled by GHerrit.

Run a preregistered 150-report forward evaluation: ten modes, three frozen
conditions, and five fresh replicates per cell, with two blind scorers per mode
and adjudication before unblinding. V2 versus V1 is the primary comparison; the
V1 core ablation is only a historical bridge.

V2 passes every whole-execution, exact Rust-1.79/1.80 boundary,
producer-quantifier, ticket, configuration, and published-contract atom. It
produces no proposal laundering and retains strong reconstructed-proof
behavior.

The release gate still fails with 16 atom misses and five hard errors. Four of
five V2 reports contract an inclusive stable-release interval by omitting Rust
1.80.1, then assert exhaustive closure. Another report assembles every fact
needed for a valid empty-slice UB witness but dilutes the conclusion to
UNPROVED by continuing to seek a universal positive lemma. Sparse-version
interval claims cause two more misses; one omitted alias route exposes an
oracle-granularity issue rather than a clear skill defect.

The evidence shows that recovering the quantified domain must itself be a
proof obligation and that verdicts need explicit logical certificates. It
motivates V3's Required/Covered model, domain-transformation obligations,
multi-release proof bases, and existential UB certificate.

Preserve the failed gate unchanged. Differences between coherent conditions
are mixed, modes are heterogeneous, five replicates are an engineering screen,
and procedural isolation and unavailable model/seed identity preclude a broad
causal or population-level claim.

gherrit-pr-id: Gthyz3viupsc7cxzrbqaql6qitqmx2ews
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 91.85%. Comparing base (df5c1ef) to head (61a4fac).

Additional details and impacted files
@@                        Coverage Diff                         @@
##           Gxw7ewqzcrigbrgikdrjkmx2nnotda4uz    #3523   +/-   ##
==================================================================
  Coverage                              91.85%   91.85%           
==================================================================
  Files                                     20       20           
  Lines                                   6093     6093           
==================================================================
  Hits                                    5597     5597           
  Misses                                   496      496           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 61a4facf53

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


## Scoring basis

The published interval is not limited to the four `.0` releases. Rust 1.80.1 was a stable release between 1.80.0 and 1.81.0 (official Rust release notes, `Version 1.80.1 (2024-08-08)`), and the exact 1.80.1 Reference/std pages exist. Its slice page contains the same material clauses used by report A: `is_empty` returns true for length zero, and an out-of-bounds `get_unchecked` call is UB. Therefore the literal policy union has 19 configurations:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Rerun the scorer that used a prohibited source

In the v2-forward blind-scoring workflow, D-s1 explicitly relies on official Rust release notes to establish the existence and date of Rust 1.80.1, even though scoring-prompt.md permits inspection only of the packet and exact-version Rust Reference or standard-library documentation. The ledger invalidates and reruns T-s2 for the same release-blog violation but incorrectly counts D-s1 among the 20 valid scores; because this score was copied into the canonical D matrix and feeds the aggregate, D-s1 must be excluded and freshly rerun before the scoring and gate artifacts can be considered protocol-valid.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants