AI-Generated Summary of Evaluator Reasoning at Dataset Level #10859
jannikmaierhoefer
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
During prompt experiments, evaluator reasoning is available per trace, but the dataset view only shows aggregated scores. This makes it difficult to understand why scores turned out the way they did or what common issues appeared without manually reviewing each trace.
Requested improvement:
Provide an AI-generated summary of evaluator reasoning at the dataset level, highlighting recurring problems, patterns, and explanations behind the aggregated scores. This would also be highly valuable for online user traces.
Additional information
No response
All reactions