v1.16.0
What's Changed
- Add deterministic agentic evaluation harness by @natiixnt in #240
- Add layer-1 evaluation results and phrasing distinguishability metric by @natiixnt in #241
- Add MCP stdio determinism end-to-end check by @natiixnt in #242
- Add agent-in-the-loop arm harness (layer 2) by @natiixnt in #243
- Add night-1 agent-arm pilot results and 56c7 failure audit by @natiixnt in #244
- Add context-heavy task corpus from django and sympy by @natiixnt in #245
- Add night-2 harness: three arms, stream-json capture, tool-call metric by @natiixnt in #246
- Add night-2 agent-arm results, arms, and coverage sweep by @natiixnt in #247
- Add the agentic-evaluation writeup by @natiixnt in #248
- Write the redcon instruction block to CLAUDE.md when it is missing by @natiixnt in #249
- Slim the MCP tool descriptions and cap their size by @natiixnt in #250
- Scale the default pack budget by repository size by @natiixnt in #251
- Release 1.16.0 by @natiixnt in #252
Full Changelog: v1.15.0...v1.16.0