attacklm-dataset v0.8.0 — held-out NLL evaluation
Sprint 4: NLL-based held-out eval suite. Implements MAI-Thinking-1 §2.3 weighted Eq-3 aggregate. New: split_held_out.py, held_out_nll.py, docs/HELD_OUT_NLL.md. 30 new tests.
Sprint 4: NLL-based held-out eval suite. Implements MAI-Thinking-1 §2.3 weighted Eq-3 aggregate. New: split_held_out.py, held_out_nll.py, docs/HELD_OUT_NLL.md. 30 new tests.