corpus: harvest realized-outcome verifier cases - #2843
Conversation
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
Workflow source neededPR #2843 needs either a linked GitHub issue or one valid non-issue Workflow Source before PR metadata automation can manage it safely. Please do one of:
Once a valid source is present, this warning will not be reposted. |
📝 WalkthroughWalkthroughThe verifier corpus staging configuration gains four harvested cases. The pilot configuration receives a ChangesVerifier corpus data
Estimated code review effort: 2 (Simple) | ~10 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Automated Status SummaryHead SHA: 0cb0a97
Coverage Overview
Updated automatically; will refresh on subsequent CI/Docker completions. Keepalive checklistScopeNo scope information available Tasks
Acceptance criteria
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ed96593ec1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", |
There was a problem hiding this comment.
Remove the mislabeled clean-pass case
PR #2528 did not remain a clean, durably complete PASS: its merge commit d096fa2 was followed only two hours later by the direct child commit 7648a7e, titled Follow up #2528: upload LangSmith fleet artifact under registry name. Because the pilot evaluates only the original PR diff, a verifier that correctly detects that missing work will return NON_PASS and be charged a false fail; with the current 47 expected-PASS cases, that single error produces a Wilson upper bound of about 0.111, exceeding the policy's 0.10 false-fail gate. Exclude this case or adjudicate it as NON_PASS rather than promoting it as clean-pass.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@config/model_eval_pilot.json`:
- Around line 217-405: Update the model evaluation configuration’s purpose text
to distinguish the frozen 30-case paired pilot from the additional harvested
cases, and revise the evaluation corpus handling so repeated PASS-only
harvesting does not dilute the NON_PASS signal used to compare model candidates.
Preserve the owner-sourced NON_PASS semantics while making the
frozen-versus-augmented case sets explicit.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 022d56a5-824a-43cd-a099-4a4bb2ef10e8
📒 Files selected for processing (2)
config/model_eval_corpus_staging.jsonconfig/model_eval_pilot.json
| { | ||
| "case_id": "workflows-2545", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2545, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2544", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2544, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2543", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2543, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2540", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2540, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2539", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2539, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2538", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2538, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2537", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2537, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2536", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2536, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2535", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2535, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2534", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2534, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2533", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2533, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2532", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2532, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2531", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2531, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2530", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2530, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2528", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2528, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2527", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2527, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2524", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2524, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2523", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2523, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2522", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2522, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2520", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2520, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| }, | ||
| { | ||
| "case_id": "workflows-2519", | ||
| "repo": "stranske/Workflows", | ||
| "pr": 2519, | ||
| "expected_verdict": "PASS", | ||
| "category": "clean-pass", | ||
| "provenance": "harvested", | ||
| "harvested_at": "2026-07-27" | ||
| } |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🔵 Trivial
Growing PASS-only harvest may dilute NON_PASS signal; "purpose" text is now stale.
All 21 newly harvested cases are PASS/clean-pass (by design, since NON_PASS semantics stay owner-sourced). This is expected, but it shifts the pilot's NON_PASS ratio from 4/30 (~13%) toward 4/51 (~8%) as harvesting continues each round, which could weaken the pilot's power to discriminate model candidates on NON_PASS detection over time. Separately, purpose (Line 4) still describes this as a "Paired 30-case pilot" even though the corpus now totals 51 cases — worth updating or clarifying that harvested cases augment, but aren't part of, the frozen 30-case paired set.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@config/model_eval_pilot.json` around lines 217 - 405, Update the model
evaluation configuration’s purpose text to distinguish the frozen 30-case paired
pilot from the additional harvested cases, and revise the evaluation corpus
handling so repeated PASS-only harvesting does not dilute the NON_PASS signal
used to compare model candidates. Preserve the owner-sourced NON_PASS semantics
while making the frozen-versus-augmented case sets explicit.
Automated corpus growth (#2819 move 2).
High-confidence cases derived from realized PR outcomes — see the run
summary for the promoted case list. Ambiguous cases were routed to the
auto-expiring staging file, not here.
Expected verdicts here come from what the world already adjudicated by
merging or reverting each PR. The semantic NON_PASS categories
(stale-verifier-claim, review-thread-debt, missing-acceptance-criterion)
remain owner-sourced and are never machine-added.
Summary by CodeRabbit