[solution · EX-19] I hand-labelled 200 answers and my judge scored κ 0.31 #155
Replies: 2 comments 2 replies
This is the best submission on this exercise so far, and it is because of the part you nearly did Measuring your own self-consistency is the step nobody takes, and without it κ is The arithmetic worth adding to your write-up: with rater self-consistency around 0.88, the |
The rubric point generalises past judging and is the more transferable half. "Partially correct" is not a labelling edge case. It is an unspecified requirement. You I would go further than "write the rubric first": label 20 examples, find the disagreements with |
Uh oh!
There was an error while loading. Please reload this page.
Submission for EX-19.
What I did
Labelled 200 answers by hand before looking at the judge's verdicts, then computed agreement.
The numbers
87.5% raw agreement, κ 0.417. On the Landis–Koch bands that is "moderate", and it is nowhere
near good enough to gate a release on.
The part that took me longest
I labelled 50, got bored, and started skimming. When I re-labelled those 50 blind two days later I
disagreed with myself on 6 of them.
That is a self-consistency of 88%, which caps the κ I can possibly measure against the judge —
I am one of the two raters and I am noisy. I threw the first 50 away and redid them properly.
What I would do differently
Write the rubric before labelling, not during. My disagreements with myself were almost all
"partially correct" cases where I had not decided in advance whether partial credit existed. That
is not a judgement problem, it is a spec problem, and the judge inherits it.
The uncomfortable conclusion
My judge is not the weak link. My labels are. A κ of 0.417 against a rater who agrees with
herself 88% of the time is close to the ceiling that noise permits.
All reactions