You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
v0.25.0 — Honest E7 SWE-bench Verified end-to-end eval
Phases 1–5 (PRs #1–#10 on base d0d497c):
#1 CI paper build · #2 hbox fixes · #3 real LLMLingua-2 + LongLLMLingua E5 + τ n-expansion
#4 README/site refresh · #5 n=300 SWE-bench_Lite headline
#6/#7 v0.24: shuffled-position E5, honest operating-point sweep, LongLLMLingua fix
#8 E7 SWE-bench Verified end-to-end (headline) · #9 v0.25 PDF refresh
#10 docs/research.html E7 sync
Headline (E7, PR #8): a real agent (aider + claude-sonnet-4-6) on SWE-bench Verified
(n=50, seed 1729, official swebench harness). Full context 52.0% (26/50); distil at its
certified trunc@500 point only 16.0% (8/50), −36pp, paired McNemar p<0.001; LLMLingua-2
26.0% (13/50). Verdict: the localization decision-equivalence certificate does NOT transfer
to end-to-end task success once compression is aggressive. The certificate (the contract)
is the contribution, not the compressor. Numbers: docs/paper/results/swe_bench_verified_e2e.json.