Skip to content

attacklm-dataset v0.8.0 — held-out NLL evaluation

Choose a tag to compare

@Veedubin Veedubin released this 14 Jul 04:11
· 6 commits to main since this release

Sprint 4: NLL-based held-out eval suite. Implements MAI-Thinking-1 §2.3 weighted Eq-3 aggregate. New: split_held_out.py, held_out_nll.py, docs/HELD_OUT_NLL.md. 30 new tests.