Replies: 1 comment
|
Thanks for sharing! We'll likely start working on this towards the end of the quarter. I will keep you posted here when we do. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
The new Score Analytics with Multi-Score Comparison feature is great for analyzing traces. It would be really useful to have similar functionality for dataset runs.
This would enable comparing multiple metrics within a single run (e.g., Human annotations vs. LLM as a Judge) or comparing the same metric across different runs (e.g., GPT-4 vs. Claude, prompt v1 vs. v2). This would make dataset evaluations much more powerful for experimenting with prompts and models.
Additional information
No response
All reactions