results-0.15.1-2026-06-26
tagged this
25 Jun 18:48
Re-ran the head-to-head on the current stack (multivon-eval 0.15.1 from PyPI, DeepEval 4.0.2, RAGAS 0.4.3), gpt-4o-mini judge, temp 0, single run. - RAGAS completes again. It errored on every case in the 0.9.8 harness; ragas 0.4.3 fixed that, so this is a real three-way comparison again (4/100 cases still error, ~13x slower than the others). - ragtruth-sum n=100 headline on 0.15.1: multivon F1 0.729 (default 0.90), 0.837 best-tuned (0.95); DeepEval 0.038 / 0.609; RAGAS 0.038 / 0.812. multivon shifted from 0.744->0.729 (within the ~+/-10pp CI at n=100; single-run, gpt-4o-mini temp-0 noise). - Killed the stale "now 0.12.0 / refresh planned" snapshot note and the 0.9.8 references; badges and tag bumped to 0.15.1 / 2026-06-26. NOTE: run.py skips frameworks whose run0 file already exists, so the stale multivon/deepeval raw files had to be deleted before this re-run actually re-measured them (the first pass silently reused 0.9.8 outputs).