Kalibera and Jones's actual contribution is not run it more times, it is that variance lives at a particular level and the repetition budget should be spent at that level. A pilot run decides whether variance is between builds, between processes, or between iterations, and the budget goes there.
Headline statistic is the median with a bootstrap confidence interval. The minimum gets its own column, because the minimum is a statistic about the luckiest run and printing it as if it were typical is the most common quiet distortion in this field.
Repetitions are added on the width of a single interval, never on whether a comparison has become significant.
Done when the loop is implemented, the adaptive stopping rule is keyed on interval width, and there is a test that a comparison driven stopping rule is not reachable.
Kalibera and Jones's actual contribution is not run it more times, it is that variance lives at a particular level and the repetition budget should be spent at that level. A pilot run decides whether variance is between builds, between processes, or between iterations, and the budget goes there.
Headline statistic is the median with a bootstrap confidence interval. The minimum gets its own column, because the minimum is a statistic about the luckiest run and printing it as if it were typical is the most common quiet distortion in this field.
Repetitions are added on the width of a single interval, never on whether a comparison has become significant.
Done when the loop is implemented, the adaptive stopping rule is keyed on interval width, and there is a test that a comparison driven stopping rule is not reachable.