What models, datasets, and tasks should AgML benchmark next? #4
masonearles
started this conversation in
Models, Benchmarks & VLMs
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
AgML's leaderboard is intended to make model evaluation across agricultural datasets more transparent, comparable, and reproducible—not simply to declare a universal winner.
Current results already show why this matters: performance can vary substantially by dataset, task, label ontology, metric, prompt, and evaluation protocol. Aggregate rankings are useful, but per-dataset behavior and methodological details are essential.
What should we benchmark next?
We welcome proposals involving:
Propose or share an evaluation
Please include as much of the following as possible:
Explore the AgML leaderboard and see the result contribution format.
Preliminary results are welcome when clearly labeled. Before a result is incorporated into the public leaderboard, we will ask for enough methodological detail to make it interpretable and reproducible.
All reactions