Skip to content
Discussion options

You must be logged in to vote

Marking this, because the shape matters as much as the content.


A stronger answer, roughly as I would say it out loud:

"It depends on whether they have labels, and usually they do not — so I'd start with the first
two weeks rather than describing a finished harness.

The eval set is the work. I need questions, and for each one the gold evidence, not just
the expected answer. Without evidence labels I cannot tell a retrieval failure from a generation
failure, and I cannot detect the case that matters most: the answer is right and nothing relevant
was retrieved. That means the model knew it already and retrieval is decorative for that query.

I'd include null questions at their real base rate

Replies: 2 comments 1 reply

Comment options

You must be logged in to vote
1 reply
@akash-coded
Comment options

akash-coded Sep 1, 2026
Maintainer Author

Comment options

You must be logged in to vote
0 replies
Answer selected by akash-coded
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
cohort Cohort-facing work casebook A simulated teaching transcript, not a real exchange
1 participant