feat(grading): add diff-scope assertions#164
Merged
Merged
Conversation
slowdini
marked this pull request as ready for review
July 24, 2026 05:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
diff_scopeassertion withmax_files_touchedandmax_lines_changedthresholdsdiff-scope.jsonand preserve raw per-run values inbenchmark.json(eval, condition, run)a private environment, exclude.eval-magic-outputs, and capture baselines after staging, guard setup, and shadow preflightWhy
The grading pipeline could check correctness, transcripts, and held-out commands, but could not measure whether an agent made a focused change. This adds scope as a deterministic secondary signal while retaining correctness assertions as the primary measure.
Before, dispatch prompts constrained all writes to the framework output directory and benchmark artifacts contained no raw change-scope context. After this change, agents may work normally inside a private task environment while eval-magic measures files, added/removed byte-lines, and zero-context Myers hunks outside the hidden framework-output subtree.
Older iterations remain readable. Their benchmark omits
diff_scope; configuring a scope assertion without a captured baseline produces rebuild guidance.Closes #116
Verification
cargo fmt --checkcargo test(554 unit, 107 CLI, 94 run tests)cargo clippy --all-targets -- -D warningsgit diff --check