Repository navigation
Releases: hyunyoungnam/Bench2Agent
Release list
Bench2Agent 0.3.0
Bench2Agent 0.3.0 renames the application, Python package, CLI and MCP server
from the previous project names. The local checkout folder can keep its name.
Install using either method:
uv tool install bench2agent
bench2agentpython -m pip install bench2agent
bench2agentThe first conversation downloads the checksum-verified research snapshot.
Use bench2agent claude or bench2agent codex for an existing client login;
bench2agent mcp exposes the same six read-only benchmark tools.
Existing users can copy research data, model settings and conversations with
bench2agent migrate --from /path/to/previous/data-home. This preserves source
files and skips existing destination files. Re-register MCP under bench2agent
and restart the client. See the migration guide.
The tested data is data-20261007:
28 editions, 58,616 parsed paper-edition observations and 9,332 catalogue entries.
The renamed data download has the same bytes, checksum and snapshot identity
as the original release asset. Benchmark locations and reviewed introduction
paper links remain part of tool results; coverage limits remain explicit.
bench2agent serve retains the local browser prototype. The shared relational
database and hosted web conversation are planned in the architecture design;
this release still queries the JSON snapshot.
Validation: 82 Python tests, 3 VS Code tests, strict package metadata checks,
fresh wheel installation, migration and representative real-snapshot MCP queries.
Python 3.10+; standard-library runtime. The package contains code, not corpus
files or downloaded research papers. Earlier published packages remain available.
Wheel and source-distribution hashes are in the attached SHA256SUMS.
BenchTrend 0.2.1 — benchmark links and simpler setup
BenchTrend 0.2.1 adds benchmark locations and introducing-paper links to conversation tools, and simplifies installation and AI-client launch.
Install or update
python -m pip install --upgrade benchtrend
benchtrendFor a uv installation, use uv tool upgrade benchtrend; for a new installation, use uv tool install benchtrend.
Changes
- The first interactive launch downloads the pinned research-data snapshot automatically and checks its SHA-256. Existing installed data is reused.
benchtrend claudeandbenchtrend codexopen the selected client's conversation with BenchTrend MCP tools using the client's existing login. No persistent registration step is required.- Benchmark results include repository, Hub dataset and homepage locations separately from reviewed introducing-paper links. Locations belong to benchmark IDs, so homonyms do not borrow each other's links. Missing locations and introducing papers remain explicit.
- API-checked mappings and authored homepages are distinguished. Check dates remain unknown for the current snapshot; its export date is never used as a check date.
- Updated OpenAI defaults, project links on PyPI, and a shared user-facing README with tested English examples and simpler installation instructions.
Data and verification
This code release uses the existing data-20261007 snapshot: 28 editions, 9,332 benchmark/dataset entries, and 58,616 parsed paper-edition observations. No new data download is required for existing users. The snapshot includes usage and paper evidence, not external datasets' underlying evaluation records.
Validation: 76 Python tests; package CI on Python 3.10 and 3.14; fresh-wheel installation and MCP smoke test outside the checkout; real-snapshot checks for counts, coverage and all four combinations of location/introducing-paper availability.
The same wheel and source distribution are supplied on GitHub and PyPI, with SHA-256 checksums attached here. BenchTrend remains a research prototype; usage frequency measures adoption, not benchmark quality.
BenchTrend v0.2.0
First code release of BenchTrend: which benchmarks the field evaluates on, how that changes across conference editions, and which newly introduced benchmarks other authors adopt — counted from the papers' own sentences. Research prototype.
What is in it
benchtrend— a terminal conversation over the OpenAI or Anthropic API (your key, hidden, never written to disk) with saved sessions,/sources, and quote and figure verification before an answer is shown.benchtrend mcp— a stdio MCP server with six read-only tools (benchmark_scope,benchmark_usage,benchmark_trend,new_benchmarks,benchmark_adoption,benchmark_evidence);benchtrend mcp --connect claude/--connect codexregisters it with the client's own CLI, so that route needs no API key.benchtrend data install— installs a released data snapshot from a file or an HTTPS URL, verifying the bundle manifest and an optional SHA-256, atomically.benchtrend status— data, API keys, and whether Claude Code / Codex have the server registered and connecting.- The
bellwethercommand for the earlier browser interface and build-machine exports.
Python 3.10+ (tested on 3.10 and 3.14), standard library only; Windows works natively for the terminal and MCP routes.
Install
uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz
# or, from the wheel attached below:
uv tool install ./benchtrend-0.2.0-py3-none-any.whlThen the data (tested with this release):
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
--sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50eData release: data-20261007 — snapshot 0dd4b9bbee4b54d2, 28 editions of 10 venues, 58,616 parsed papers, 9,332 benchmarks and datasets. Code and data are versioned separately; see docs/releases.md.
What it does not say
Only stated use is counted ("we evaluate on …"), over papers with parsed arXiv full text — the denominator is printed beside every figure. "New" means first claimed inside this corpus, which starts in 2023. Frequency is never quality. Full guide: docs/terminal.md.
Checksums
SHA256SUMS is attached; verify with sha256sum -c SHA256SUMS.
Benchmark data 2026-10-07
Benchmark usage snapshot for benchtrend: 28 editions (ICML, ICLR, NeurIPS, CVPR, ICCV, ECCV, AAAI, ACL, EMNLP, CoRL; three most recent each, two for ICCV/ECCV), 58,616 parsed papers, 9,332 benchmarks with stated evaluation/training roles and verbatim evidence. Snapshot id 0dd4b9bbee4b54d2.
Install:
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e
Parsed full text (six editions)
Input for icml.resources and for paper_text. Not part of the data bundle; not tracked in git (per-file > 100 MB).
Data bundle 2026-09-23
Site + search index for six editions, now with data/processed/resources.json (10,330 papers' code/data/model links, 9,839 links with GitHub/Hub metadata) and the chat.html fixes (clickable links, one card per paper, Korean off by default). Install: bellwether fetch-data --release data-20260923
Data bundle 2026-09-09
Site, chat shell and search index for six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26).
Since data-20260903:
- the selected set has four views — cards, a table of the extracted fields, a map placed by the selection's own vectors, and the embedding subgroups
- a reading queue (read / later / not mine) that reorders nothing
- .bib and .csv for a selection; markdown or JSON for a conversation
- figures in an answer are checked by running their tool again
- a stop button, conversation addresses, and the cite-chip deep link repaired
Data bundle 2026-09-03
Six editions, 29,605 papers — now including the Bellwether chat shell, card_terms (alias-canonicalized names), and the derivation-tree era site. Unpack: bellwether fetch-data --url
Data bundle 2026-09-02
Six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26), 29,605 papers: built site + search index + processed files. Unpack with: wnai fetch-data --file wnai-data-20260902.tar.gz