Repository navigation
Releases: hyunyoungnam/BenchTrend
Release list
BenchTrend 0.2.1 — benchmark links and simpler setup
BenchTrend 0.2.1 adds benchmark locations and introducing-paper links to conversation tools, and simplifies installation and AI-client launch.
Install or update
python -m pip install --upgrade benchtrend
benchtrendFor a uv installation, use uv tool upgrade benchtrend; for a new installation, use uv tool install benchtrend.
Changes
- The first interactive launch downloads the pinned research-data snapshot automatically and checks its SHA-256. Existing installed data is reused.
benchtrend claudeandbenchtrend codexopen the selected client's conversation with BenchTrend MCP tools using the client's existing login. No persistent registration step is required.- Benchmark results include repository, Hub dataset and homepage locations separately from reviewed introducing-paper links. Locations belong to benchmark IDs, so homonyms do not borrow each other's links. Missing locations and introducing papers remain explicit.
- API-checked mappings and authored homepages are distinguished. Check dates remain unknown for the current snapshot; its export date is never used as a check date.
- Updated OpenAI defaults, project links on PyPI, and a shared user-facing README with tested English examples and simpler installation instructions.
Data and verification
This code release uses the existing data-20261007 snapshot: 28 editions, 9,332 benchmark/dataset entries, and 58,616 parsed paper-edition observations. No new data download is required for existing users. The snapshot includes usage and paper evidence, not external datasets' underlying evaluation records.
Validation: 76 Python tests; package CI on Python 3.10 and 3.14; fresh-wheel installation and MCP smoke test outside the checkout; real-snapshot checks for counts, coverage and all four combinations of location/introducing-paper availability.
The same wheel and source distribution are supplied on GitHub and PyPI, with SHA-256 checksums attached here. BenchTrend remains a research prototype; usage frequency measures adoption, not benchmark quality.
BenchTrend v0.2.0
First code release of BenchTrend: which benchmarks the field evaluates on, how that changes across conference editions, and which newly introduced benchmarks other authors adopt — counted from the papers' own sentences. Research prototype.
What is in it
benchtrend— a terminal conversation over the OpenAI or Anthropic API (your key, hidden, never written to disk) with saved sessions,/sources, and quote and figure verification before an answer is shown.benchtrend mcp— a stdio MCP server with six read-only tools (benchmark_scope,benchmark_usage,benchmark_trend,new_benchmarks,benchmark_adoption,benchmark_evidence);benchtrend mcp --connect claude/--connect codexregisters it with the client's own CLI, so that route needs no API key.benchtrend data install— installs a released data snapshot from a file or an HTTPS URL, verifying the bundle manifest and an optional SHA-256, atomically.benchtrend status— data, API keys, and whether Claude Code / Codex have the server registered and connecting.- The
bellwethercommand for the earlier browser interface and build-machine exports.
Python 3.10+ (tested on 3.10 and 3.14), standard library only; Windows works natively for the terminal and MCP routes.
Install
uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz
# or, from the wheel attached below:
uv tool install ./benchtrend-0.2.0-py3-none-any.whlThen the data (tested with this release):
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
--sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50eData release: data-20261007 — snapshot 0dd4b9bbee4b54d2, 28 editions of 10 venues, 58,616 parsed papers, 9,332 benchmarks and datasets. Code and data are versioned separately; see docs/releases.md.
What it does not say
Only stated use is counted ("we evaluate on …"), over papers with parsed arXiv full text — the denominator is printed beside every figure. "New" means first claimed inside this corpus, which starts in 2023. Frequency is never quality. Full guide: docs/terminal.md.
Checksums
SHA256SUMS is attached; verify with sha256sum -c SHA256SUMS.
Benchmark data 2026-10-07
Benchmark usage snapshot for benchtrend: 28 editions (ICML, ICLR, NeurIPS, CVPR, ICCV, ECCV, AAAI, ACL, EMNLP, CoRL; three most recent each, two for ICCV/ECCV), 58,616 parsed papers, 9,332 benchmarks with stated evaluation/training roles and verbatim evidence. Snapshot id 0dd4b9bbee4b54d2.
Install:
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e
Parsed full text (six editions)
Input for icml.resources and for paper_text. Not part of the data bundle; not tracked in git (per-file > 100 MB).
Data bundle 2026-09-23
Site + search index for six editions, now with data/processed/resources.json (10,330 papers' code/data/model links, 9,839 links with GitHub/Hub metadata) and the chat.html fixes (clickable links, one card per paper, Korean off by default). Install: bellwether fetch-data --release data-20260923
Data bundle 2026-09-09
Site, chat shell and search index for six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26).
Since data-20260903:
- the selected set has four views — cards, a table of the extracted fields, a map placed by the selection's own vectors, and the embedding subgroups
- a reading queue (read / later / not mine) that reorders nothing
- .bib and .csv for a selection; markdown or JSON for a conversation
- figures in an answer are checked by running their tool again
- a stop button, conversation addresses, and the cite-chip deep link repaired
Data bundle 2026-09-03
Six editions, 29,605 papers — now including the Bellwether chat shell, card_terms (alias-canonicalized names), and the derivation-tree era site. Unpack: bellwether fetch-data --url
Data bundle 2026-09-02
Six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26), 29,605 papers: built site + search index + processed files. Unpack with: wnai fetch-data --file wnai-data-20260902.tar.gz