[poll] Would you let an LLM judge gate a release? #140
akash-coded
started this conversation in
Polls
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Your judge agrees with human labels at κ = 0.64 on a 85/15 split, is self-consistent on 97% of
repeated inputs, and costs $0.004 per judged answer.
Do you let it block a merge?
The distinction the poll is really testing
A judge with moderate but consistent error can still rank two systems correctly, because a
systematic bias applies to both arms and cancels in a paired comparison. A judge with high but
erratic error cannot.
So "is κ high enough" is the wrong question. The right one is in The Calibration
Triangle.
How to vote. The API cannot create a native poll widget, so vote with a reaction on this post — 👍 A · 🎉 B · 😄 C · 🚀 D — and reply if your answer needs a condition attached. The reveal is posted as a reply in about a week, never by editing this body: editing would destroy the record of what you were actually asked.
All reactions