Replies: 1 comment
|
Hey, thanks for suggesting this. Agree this is a useful feature. I will keep you posted here if we start working on either. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
I’d like the ability to filter dataset runs by metadata or by tags. I know tags for dataset runs are not yet available but are planned, and this feature would complement that roadmap.
We use dataset runs in our CI as part of a sanity test process. These CI runs currently get mixed together with our real evaluation runs, which pollutes the score/latency graphs. Being able to filter runs - either by metadata or by tags - would allow us to cleanly separate CI runs from actual evaluation runs. This becomes even more important as we move towards supporting multiple dataset versions.
Additional information
No response
All reactions