Skip to content
Discussion options

You must be logged in to vote

Marking this as the answer — the model response, then the rubric.


A strong answer, roughly as I would give it:

"First — what would 'wrong' mean here? The number, the direction, or the mechanism? Those need
different experiments and only the third is interesting.

Assume the mechanism. Their claim is that better chunk boundaries improve retrieval, which
implies a precondition: their corpus had bad boundaries. So the falsifying experiment is a
corpus where boundaries are already good — structurally chunked documents carrying their heading
path, each chunk readable in isolation.

If their method still shows 4 points there, it is not doing what they say. It may be doing
something valuable, bu…

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by akash-coded
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
area: evaluation Metrics, judge, release gate casebook A simulated teaching transcript, not a real exchange interview-round A full simulated interview loop
1 participant