Do models catch flawed ML results? Flawed reports with byte-identical matched controls.
-
Updated
Aug 20, 2026 - Python
Do models catch flawed ML results? Flawed reports with byte-identical matched controls.
Inverted generation: construct the label first, render the sentence second. Strictly-labelled, contamination-free synthetic data without a judge.
Predict how a policy games a reward. RL environments scored by an executable-exploit verifier.
ML training defects that never crash. An agentic benchmark with patch-and-rerun, unit-test verification.
A live, contamination-resistant benchmark for evaluating LLM agents on U.S. macroeconomic nowcasting.
To associate your repository with the contamination-free topic, visit your repo's landing page and select "manage topics."