Skip to content

Repository files navigation

BenchPress

Evaluation Planning with Cognitive-Ability Tags Aligned to LLM Benchmark Score Patterns

https://ssu-nlp.github.io/BenchPress/

Demo video

License: MIT

BenchPress system overview: Leaderboard Scores and Benchmark Corpus feed the Autotagging Loop (read benchmark item, tag cognitive abilities, build benchmark profile), which produces neighbor benchmarks, a customized benchmark set, and cognitive ability tags.

How it works

The system is a closed alignment loop with four LLM roles, driven by a single objective: the cosine similarity between two benchmarks' ability-tag profiles should track how similarly the two benchmarks rank models.

  • Mapper — a lightweight model that tags each benchmark item on every ability axis. It runs at high concurrency and produces the raw per-item labels that aggregate into ability profiles.
  • Executer — regenerates the ability vocabulary each iteration, proposing candidate axis sets of varying size.
  • Maker — reduces per-item evidence into stable per-benchmark profiles via a cached map-reduce pass.
  • Improver — proposes revisions to the tagging prompt and vocabulary. Samples are accepted only when they pass an alignment-improvement gate; otherwise the previous state is kept.

The loop optimizes an alignment loss (L_align, the error between tag-similarity and score-ranking-similarity over benchmark pairs) alongside rank-correlation diagnostics (ρ_align) and a tagging-quality term (Δ_tag). A candidate survives an iteration only if it improves these metrics, so the vocabulary is grounded in observed score behavior rather than in human intuition about categories.

Two stages produce the final artifacts:

  • Pre-experiment (Part 1)autotagging_loop/pretrain.py runs a single global alignment loop and emits the seed taxonomy: final/I_star.txt plus data/cognitive_abilities.json (the seed ability vocabulary).
  • Main experiment (Part 2)autotagging_loop/main.py (via the runner) reuses those seed artifacts and runs the validated loop over held-out splits of benchmarks and models, with best-iteration selection tuned for stability and generalization.

Quickstart

Requires Python ≥ 3.10 and uv.

git clone https://github.com/SSU-NLP/BenchPress.git
cd BenchPress
uv sync
cp .env.example .env    # fill in your API keys

The .env file configures two provider endpoints: a small model for the Mapper and a larger model shared by the Executer, Maker, and Improver. Each role in benchpress_config.json names its own base_url_env / api_key_env, so roles can point at different providers without any code changes.

Running the demo locally

The demo has two parts. The Composer is the publishing backend; the Builder is the frontend.

# Publishing requires Hugging Face credentials (skip if you only want to preview)
uv run hf auth login --token "$HF_TOKEN"

# Composer — publishing backend → http://127.0.0.1:7860
uv run python benchpress/space/app.py

# Builder — frontend → http://localhost:5173/BenchPress/
cd benchpress/benchboard && npm install && npm run dev

The Composer requires no account, login, or user-provided token to preview compositions. Publishing does require Hugging Face credentials: uv run hf auth login login once before starting app.py and the Composer will pick up your credentials automatically. Generated demo repositories are public and prefixed with demo-. See benchpress/space/README.md for the relevant Space secrets and deployment steps.

Citation

@misc{jang2026benchpress,
  title  = {BenchPress: Capability-Targeted Evaluation Planning from LLM Benchmark Score Patterns},
  author = {Jang, Giyoon and Cho, Seonghyeon and Kim, Doyun and Song, Minsoo and
            Song, Jihun and Choo, Kyojun and Nam, Chailin and Park, Chanjun},
  year   = {2026},
  url    = {https://github.com/SSU-NLP/BenchPress}
}

License

Released under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages