[lab · T5] Judgement — track thread #121
Unanswered
akash-coded
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What this track is about
94% agreement sounds like a working judge. Work the arithmetic and most of it was available for free.
Where people get stuck
Run the class-balance check and watch the verdict move two bands on the same judge and the same raw agreement. Then remember the conclusion that matters: a judge with moderate but consistent error can still rank two systems correctly; a high but erratic one cannot.
How to ask here
Paste the check that failed and its message. The messages are written to be actionable — if
one is not, that is a bug in the lab and saying so is a useful contribution.
Submitting
Fork, fill in
starter.py, open a pull request. TheLabs workflow posts your results as a comment and
edits it in place on each push. It will not fail the build on red checks — the comment says
what failed.
What this track hands to the next one
Quality produces the artefact the following stage needs. If you finish L09 and cannot say what
that artefact is for your own corpus, the track is not done — that is the point of the lifecycle
running underneath the pathway rather than beside it.
All reactions