We introduce the Self-Challenge framework. As shown in Self-Challenge framework provides an evaluation protocol that asks LLMs to find their own limitations from the errors they make, with the help of human-in-the-loop.
Applying this framework to GPT-4, we discover 8 error patterns from 189 instances and then build a benchmark, SC-G4, consisting of 1,835 instances generated by GPT-4 using these patterns, with human-annotated gold responses.
You can directly download the data from this fold.
This work is accepted by COLM 2024. If you find our paper interesting, please kindly cite our paper.
