Skip to content
Discussion options

You must be logged in to vote

Both verdicts are correct, and understanding why is worth more than the answer to "should I
ship".

Why the intervals differ despite equal point estimates. evidence_recall is continuous —
a question can score 0, 0.5 or 1.0 depending on how many gold items arrived.
full_chain_recall is binary per question: 1 only if every gold item arrived, 0
otherwise. Binary outcomes have much higher variance per sample, so the same effect size needs
more questions to become distinguishable from zero.

You have not measured two different effects. You have measured one effect with two different
amounts of statistical power.

So do you ship it? On this evidence, yes — with the honest sentence attached:

"Ship…

Replies: 4 comments 6 replies

Comment options

You must be logged in to vote
4 replies
@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

Answer selected by akash-coded
Comment options

You must be logged in to vote
2 replies
@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

@akash-coded
Comment options

akash-coded Aug 31, 2026
Maintainer Author

Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
casebook A simulated teaching transcript, not a real exchange
1 participant