Replies: 1 comment
|
Hey, thanks for sharing feedback. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
I am requesting a feature to enhance the evaluation of non-deterministic systems, particularly those with LLMs. Due to the inherent variability of these models, a single-run evaluation score is insufficient for accurately assessing system reliability.
To provide a statistically robust assessment of system performance and stability, I propose adding support for calculating and tracking aggregated statistical measures over multiple evaluation runs against the same test set.
Core Requirements
Statistical Metrics: Calculate and display the following metrics for evaluation scores across a series of runs:
Aggregation Grouping: Allow users to explicitly group multiple evaluation runs so that the system knows which scores to aggregate.
Visualization: Present the score distributions clearly in the UI using histograms or box plots alongside the numerical statistics. This will help users visualize the variance and shape of the score distribution.
Benefits
This capability is crucial for:
Quantifying the uncertainty in a system's performance.
Establishing a reliable baseline for performance metrics.
Tracking system stability and variance across different versions (e.g., before and after a model fine-tune) and making the reported evaluation scores statistically meaningful.
Additional information
No response
All reactions