My reranker improved evidence recall but full-chain recall is 'inside the noise band'. Do I ship it? #29
The question in one line
What I have already triedRe-ran both. Same verdicts. Checked I am comparing the same question set. The numbersThe point estimates are nearly identical. Why do the verdicts differ? WhereNotebook 04 §4.10 |
Replies: 4 comments 6 replies
|
Both verdicts are correct, and understanding why is worth more than the answer to "should I Why the intervals differ despite equal point estimates. You have not measured two different effects. You have measured one effect with two different So do you ship it? On this evidence, yes — with the honest sentence attached:
That is close to verbatim what the capstone's decision record generates, and it is the sentence Two things to do next, in this order:
The thing not to do is pick whichever metric cleared and report only that one. Everyone can Seeded by faculty as a worked example of the format. Start your own thread rather than replying here unless you have something to add. |
On the practical side — can you not just push the bootstrap harder? If the confidence interval on full-chain is I have done this before on a different project and the intervals definitely got narrower. |
Since the question underneath all this is "how many questions would I actually need", here is the arithmetic rather than a shrug. This is the part interviewers push on and most candidates hand-wave. What you have. 207 questions, observed delta For the interval to just exclude zero, the half-width has to fall below the point estimate. Half-width scales as But 383 is the optimistic answer, and you should not quote it. It assumes the true effect equals the one you measured, and a delta that only just misses significance is more likely than not an overestimate — the winner's curse. Budget roughly double: The thing worth internalising: on a binary metric like full-chain recall, per-question variance is Practical consequence: growing the eval set from 207 to ~800 is a day of |
Following up three weeks later, since this thread taught me more than the change did. I did what Dana said and grew the frozen slice to 812 questions before touching the reranker again. Re-ran the identical change: Two things I did not expect.
Shipped on evidence recall with the full-chain result stated. The reviewer's question was "how do you know 207 was too few" and I had an actual answer, which I would not have had before this thread. |
Both verdicts are correct, and understanding why is worth more than the answer to "should I
ship".
Why the intervals differ despite equal point estimates.
evidence_recallis continuous —a question can score 0, 0.5 or 1.0 depending on how many gold items arrived.
full_chain_recallis binary per question: 1 only if every gold item arrived, 0otherwise. Binary outcomes have much higher variance per sample, so the same effect size needs
more questions to become distinguishable from zero.
You have not measured two different effects. You have measured one effect with two different
amounts of statistical power.
So do you ship it? On this evidence, yes — with the honest sentence attached: