Skip to content
Discussion options

You must be logged in to vote

Smaller than most people expect for catching regressions, larger than most expect for ranking two close variants. Twenty carefully chosen examples covering the failure modes you actually care about will reliably catch a regression when a model version changes. Distinguishing two prompts that differ by a few percent needs enough examples that the difference exceeds sampling noise, which is usually in the low hundreds.

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by martex-dev
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
1 participant