Skip to content

Releases: hyunyoungnam/BenchTrend

BenchTrend 0.2.1 — benchmark links and simpler setup

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 05:25

BenchTrend 0.2.1 adds benchmark locations and introducing-paper links to conversation tools, and simplifies installation and AI-client launch.

Install or update

python -m pip install --upgrade benchtrend
benchtrend

For a uv installation, use uv tool upgrade benchtrend; for a new installation, use uv tool install benchtrend.

Changes

  • The first interactive launch downloads the pinned research-data snapshot automatically and checks its SHA-256. Existing installed data is reused.
  • benchtrend claude and benchtrend codex open the selected client's conversation with BenchTrend MCP tools using the client's existing login. No persistent registration step is required.
  • Benchmark results include repository, Hub dataset and homepage locations separately from reviewed introducing-paper links. Locations belong to benchmark IDs, so homonyms do not borrow each other's links. Missing locations and introducing papers remain explicit.
  • API-checked mappings and authored homepages are distinguished. Check dates remain unknown for the current snapshot; its export date is never used as a check date.
  • Updated OpenAI defaults, project links on PyPI, and a shared user-facing README with tested English examples and simpler installation instructions.

Data and verification

This code release uses the existing data-20261007 snapshot: 28 editions, 9,332 benchmark/dataset entries, and 58,616 parsed paper-edition observations. No new data download is required for existing users. The snapshot includes usage and paper evidence, not external datasets' underlying evaluation records.

Validation: 76 Python tests; package CI on Python 3.10 and 3.14; fresh-wheel installation and MCP smoke test outside the checkout; real-snapshot checks for counts, coverage and all four combinations of location/introducing-paper availability.

The same wheel and source distribution are supplied on GitHub and PyPI, with SHA-256 checksums attached here. BenchTrend remains a research prototype; usage frequency measures adoption, not benchmark quality.

BenchTrend v0.2.0

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 02:48

First code release of BenchTrend: which benchmarks the field evaluates on, how that changes across conference editions, and which newly introduced benchmarks other authors adopt — counted from the papers' own sentences. Research prototype.

What is in it

  • benchtrend — a terminal conversation over the OpenAI or Anthropic API (your key, hidden, never written to disk) with saved sessions, /sources, and quote and figure verification before an answer is shown.
  • benchtrend mcp — a stdio MCP server with six read-only tools (benchmark_scope, benchmark_usage, benchmark_trend, new_benchmarks, benchmark_adoption, benchmark_evidence); benchtrend mcp --connect claude / --connect codex registers it with the client's own CLI, so that route needs no API key.
  • benchtrend data install — installs a released data snapshot from a file or an HTTPS URL, verifying the bundle manifest and an optional SHA-256, atomically.
  • benchtrend status — data, API keys, and whether Claude Code / Codex have the server registered and connecting.
  • The bellwether command for the earlier browser interface and build-machine exports.

Python 3.10+ (tested on 3.10 and 3.14), standard library only; Windows works natively for the terminal and MCP routes.

Install

uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz
# or, from the wheel attached below:
uv tool install ./benchtrend-0.2.0-py3-none-any.whl

Then the data (tested with this release):

benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
  --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e

Data release: data-20261007 — snapshot 0dd4b9bbee4b54d2, 28 editions of 10 venues, 58,616 parsed papers, 9,332 benchmarks and datasets. Code and data are versioned separately; see docs/releases.md.

What it does not say

Only stated use is counted ("we evaluate on …"), over papers with parsed arXiv full text — the denominator is printed beside every figure. "New" means first claimed inside this corpus, which starts in 2023. Frequency is never quality. Full guide: docs/terminal.md.

Checksums

SHA256SUMS is attached; verify with sha256sum -c SHA256SUMS.

Benchmark data 2026-10-07

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 01:32

Benchmark usage snapshot for benchtrend: 28 editions (ICML, ICLR, NeurIPS, CVPR, ICCV, ECCV, AAAI, ACL, EMNLP, CoRL; three most recent each, two for ICCV/ECCV), 58,616 parsed papers, 9,332 benchmarks with stated evaluation/training roles and verbatim evidence. Snapshot id 0dd4b9bbee4b54d2.

Install:

benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e

Parsed full text (six editions)

Choose a tag to compare

Input for icml.resources and for paper_text. Not part of the data bundle; not tracked in git (per-file > 100 MB).

Data bundle 2026-09-23

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 23 Sep 02:41

Site + search index for six editions, now with data/processed/resources.json (10,330 papers' code/data/model links, 9,839 links with GitHub/Hub metadata) and the chat.html fixes (clickable links, one card per paper, Korean off by default). Install: bellwether fetch-data --release data-20260923

Data bundle 2026-09-09

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 09 Sep 07:34

Site, chat shell and search index for six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26).

Since data-20260903:

  • the selected set has four views — cards, a table of the extracted fields, a map placed by the selection's own vectors, and the embedding subgroups
  • a reading queue (read / later / not mine) that reorders nothing
  • .bib and .csv for a selection; markdown or JSON for a conversation
  • figures in an answer are checked by running their tool again
  • a stop button, conversation addresses, and the cite-chip deep link repaired

Data bundle 2026-09-03

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 03 Sep 08:06

Six editions, 29,605 papers — now including the Bellwether chat shell, card_terms (alias-canonicalized names), and the derivation-tree era site. Unpack: bellwether fetch-data --url

Data bundle 2026-09-02

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 02 Sep 07:42

Six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26), 29,605 papers: built site + search index + processed files. Unpack with: wnai fetch-data --file wnai-data-20260902.tar.gz