Repository navigation
BenchTrend 0.2.1 — benchmark links and simpler setup
BenchTrend 0.2.1 adds benchmark locations and introducing-paper links to conversation tools, and simplifies installation and AI-client launch.
Install or update
python -m pip install --upgrade benchtrend
benchtrendFor a uv installation, use uv tool upgrade benchtrend; for a new installation, use uv tool install benchtrend.
Changes
- The first interactive launch downloads the pinned research-data snapshot automatically and checks its SHA-256. Existing installed data is reused.
benchtrend claudeandbenchtrend codexopen the selected client's conversation with BenchTrend MCP tools using the client's existing login. No persistent registration step is required.- Benchmark results include repository, Hub dataset and homepage locations separately from reviewed introducing-paper links. Locations belong to benchmark IDs, so homonyms do not borrow each other's links. Missing locations and introducing papers remain explicit.
- API-checked mappings and authored homepages are distinguished. Check dates remain unknown for the current snapshot; its export date is never used as a check date.
- Updated OpenAI defaults, project links on PyPI, and a shared user-facing README with tested English examples and simpler installation instructions.
Data and verification
This code release uses the existing data-20261007 snapshot: 28 editions, 9,332 benchmark/dataset entries, and 58,616 parsed paper-edition observations. No new data download is required for existing users. The snapshot includes usage and paper evidence, not external datasets' underlying evaluation records.
Validation: 76 Python tests; package CI on Python 3.10 and 3.14; fresh-wheel installation and MCP smoke test outside the checkout; real-snapshot checks for counts, coverage and all four combinations of location/introducing-paper availability.
The same wheel and source distribution are supplied on GitHub and PyPI, with SHA-256 checksums attached here. BenchTrend remains a research prototype; usage frequency measures adoption, not benchmark quality.