Skip to content

Releases: hyunyoungnam/Bench2Agent

Bench2Agent 0.3.0

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 08 Oct 06:00

Bench2Agent 0.3.0 renames the application, Python package, CLI and MCP server
from the previous project names. The local checkout folder can keep its name.

Install using either method:

uv tool install bench2agent
bench2agent
python -m pip install bench2agent
bench2agent

The first conversation downloads the checksum-verified research snapshot.
Use bench2agent claude or bench2agent codex for an existing client login;
bench2agent mcp exposes the same six read-only benchmark tools.

Existing users can copy research data, model settings and conversations with
bench2agent migrate --from /path/to/previous/data-home. This preserves source
files and skips existing destination files. Re-register MCP under bench2agent
and restart the client. See the migration guide.

The tested data is data-20261007:
28 editions, 58,616 parsed paper-edition observations and 9,332 catalogue entries.
The renamed data download has the same bytes, checksum and snapshot identity
as the original release asset. Benchmark locations and reviewed introduction
paper links remain part of tool results; coverage limits remain explicit.

bench2agent serve retains the local browser prototype. The shared relational
database and hosted web conversation are planned in the architecture design;
this release still queries the JSON snapshot.

Validation: 82 Python tests, 3 VS Code tests, strict package metadata checks,
fresh wheel installation, migration and representative real-snapshot MCP queries.
Python 3.10+; standard-library runtime. The package contains code, not corpus
files or downloaded research papers. Earlier published packages remain available.

Wheel and source-distribution hashes are in the attached SHA256SUMS.

BenchTrend 0.2.1 — benchmark links and simpler setup

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 05:25

BenchTrend 0.2.1 adds benchmark locations and introducing-paper links to conversation tools, and simplifies installation and AI-client launch.

Install or update

python -m pip install --upgrade benchtrend
benchtrend

For a uv installation, use uv tool upgrade benchtrend; for a new installation, use uv tool install benchtrend.

Changes

  • The first interactive launch downloads the pinned research-data snapshot automatically and checks its SHA-256. Existing installed data is reused.
  • benchtrend claude and benchtrend codex open the selected client's conversation with BenchTrend MCP tools using the client's existing login. No persistent registration step is required.
  • Benchmark results include repository, Hub dataset and homepage locations separately from reviewed introducing-paper links. Locations belong to benchmark IDs, so homonyms do not borrow each other's links. Missing locations and introducing papers remain explicit.
  • API-checked mappings and authored homepages are distinguished. Check dates remain unknown for the current snapshot; its export date is never used as a check date.
  • Updated OpenAI defaults, project links on PyPI, and a shared user-facing README with tested English examples and simpler installation instructions.

Data and verification

This code release uses the existing data-20261007 snapshot: 28 editions, 9,332 benchmark/dataset entries, and 58,616 parsed paper-edition observations. No new data download is required for existing users. The snapshot includes usage and paper evidence, not external datasets' underlying evaluation records.

Validation: 76 Python tests; package CI on Python 3.10 and 3.14; fresh-wheel installation and MCP smoke test outside the checkout; real-snapshot checks for counts, coverage and all four combinations of location/introducing-paper availability.

The same wheel and source distribution are supplied on GitHub and PyPI, with SHA-256 checksums attached here. BenchTrend remains a research prototype; usage frequency measures adoption, not benchmark quality.

BenchTrend v0.2.0

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 02:48

First code release of BenchTrend: which benchmarks the field evaluates on, how that changes across conference editions, and which newly introduced benchmarks other authors adopt — counted from the papers' own sentences. Research prototype.

What is in it

  • benchtrend — a terminal conversation over the OpenAI or Anthropic API (your key, hidden, never written to disk) with saved sessions, /sources, and quote and figure verification before an answer is shown.
  • benchtrend mcp — a stdio MCP server with six read-only tools (benchmark_scope, benchmark_usage, benchmark_trend, new_benchmarks, benchmark_adoption, benchmark_evidence); benchtrend mcp --connect claude / --connect codex registers it with the client's own CLI, so that route needs no API key.
  • benchtrend data install — installs a released data snapshot from a file or an HTTPS URL, verifying the bundle manifest and an optional SHA-256, atomically.
  • benchtrend status — data, API keys, and whether Claude Code / Codex have the server registered and connecting.
  • The bellwether command for the earlier browser interface and build-machine exports.

Python 3.10+ (tested on 3.10 and 3.14), standard library only; Windows works natively for the terminal and MCP routes.

Install

uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz
# or, from the wheel attached below:
uv tool install ./benchtrend-0.2.0-py3-none-any.whl

Then the data (tested with this release):

benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
  --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e

Data release: data-20261007 — snapshot 0dd4b9bbee4b54d2, 28 editions of 10 venues, 58,616 parsed papers, 9,332 benchmarks and datasets. Code and data are versioned separately; see docs/releases.md.

What it does not say

Only stated use is counted ("we evaluate on …"), over papers with parsed arXiv full text — the denominator is printed beside every figure. "New" means first claimed inside this corpus, which starts in 2023. Frequency is never quality. Full guide: docs/terminal.md.

Checksums

SHA256SUMS is attached; verify with sha256sum -c SHA256SUMS.

Benchmark data 2026-10-07

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 07 Oct 01:32

Benchmark usage snapshot for benchtrend: 28 editions (ICML, ICLR, NeurIPS, CVPR, ICCV, ECCV, AAAI, ACL, EMNLP, CoRL; three most recent each, two for ICCV/ECCV), 58,616 parsed papers, 9,332 benchmarks with stated evaluation/training roles and verbatim evidence. Snapshot id 0dd4b9bbee4b54d2.

Install:

benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e

Parsed full text (six editions)

Choose a tag to compare

Input for icml.resources and for paper_text. Not part of the data bundle; not tracked in git (per-file > 100 MB).

Data bundle 2026-09-23

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 23 Sep 02:41

Site + search index for six editions, now with data/processed/resources.json (10,330 papers' code/data/model links, 9,839 links with GitHub/Hub metadata) and the chat.html fixes (clickable links, one card per paper, Korean off by default). Install: bellwether fetch-data --release data-20260923

Data bundle 2026-09-09

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 09 Sep 07:34

Site, chat shell and search index for six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26).

Since data-20260903:

  • the selected set has four views — cards, a table of the extracted fields, a map placed by the selection's own vectors, and the embedding subgroups
  • a reading queue (read / later / not mine) that reorders nothing
  • .bib and .csv for a selection; markdown or JSON for a conversation
  • figures in an answer are checked by running their tool again
  • a stop button, conversation addresses, and the cite-chip deep link repaired

Data bundle 2026-09-03

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 03 Sep 08:06

Six editions, 29,605 papers — now including the Bellwether chat shell, card_terms (alias-canonicalized names), and the derivation-tree era site. Unpack: bellwether fetch-data --url

Data bundle 2026-09-02

Choose a tag to compare

@hyunyoungnam hyunyoungnam released this 02 Sep 07:42

Six editions (ICML 2025/26, NeurIPS 2024/25, ICLR 2025/26), 29,605 papers: built site + search index + processed files. Unpack with: wnai fetch-data --file wnai-data-20260902.tar.gz