Certified reference models for AI evals: deterministically broken in one documented way each, so a benchmark's detection rate is measurable against ground truth.
-
Updated
Aug 16, 2026 - Python
Certified reference models for AI evals: deterministically broken in one documented way each, so a benchmark's detection rate is measurable against ground truth.
Add a description, image, and links to the defect-injection topic page so that developers can more easily learn about it.
To associate your repository with the defect-injection topic, visit your repo's landing page and select "manage topics."