[clinic · EX-04] Manufacture an eval set for a new domain #65
Unanswered
akash-coded
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What this exercise asks
Read the exercise first; this thread is for when you are stuck, not instead of it.
Three ways people get stuck on this one
1. Writing the questions first and looking for evidence afterwards. That reintroduces exactly the annotation-error floor the fact-graph construction exists to remove.
2. Generating only questions the retriever can already answer. An eval set with no failures cannot measure an improvement.
3. No null questions. Without them you cannot measure abstention at all, and the set silently rewards a system that never refuses.
How to ask here
Post the numbers you got, the config that produced them, and what you already tried.
A question with a measurement in it gets answered in one round trip; a question without one
takes three.
If you have finished it, post the write-up in Show and Tell with the
[solution · EX-04]prefix rather than here — this thread stays for people who are stuck.All reactions