QuantumBench is an LLM benchmark built from 769 multiple-choice questions curated from open quantum science and engineering course materials. The dataset aggregates content from MIT OCW, TUD OCW, and LibreTexts. Each question includes seven distractors, the correct answer, source metadata, and human annotations for difficulty and required expertise. The repository ships both the dataset (as a password-protected archive) and a reference evaluation script that targets OpenAI-compatible APIs.
⚠️ Important:quantumbench.zipis password-protected. Unlock it withdo_not_use_quantumbench_for_training—the password is also a reminder to keep the dataset for evaluation rather than model training.
This repository now includes a complete benchmarking agent for Qiskit Code Assistant with minimal configuration:
Quick Start:
export OPENAI_API_KEY="your_ibm_cloud_api_key"
python code/qiskit_benchmark_agent.py --analyzeFeatures:
- ✅ One-command execution - Just set your API key and run
- ✅ Automatic analysis - Detailed reports by difficulty, expertise, and subdomain
- ✅ Prompt comparison - Compare zero-shot vs chain-of-thought reasoning
- ✅ GitHub Actions - Automated benchmarking with private results
- ✅ Sensible defaults - Pre-configured for Qiskit Code Assistant endpoints
Documentation:
- Quick Start Guide - Get running in 5 minutes
- Full Documentation - Complete usage guide
- Troubleshooting Guide - Fix common issues (404 errors, auth, etc.)
- Comparison Guide - Optimize prompt types
- GitHub Actions Setup - Automated workflows
Example Output: The agent generates comprehensive analysis including pass rates by difficulty level (1-5), expertise level (1-4), subdomain (Quantum Mechanics, Computation, etc.), and question type (Algebraic, Numerical, Conceptual).
quantumbench.zip: Password-protected archive that expands into thequantumbench/directory when unlocked.quantumbench/quantumbench.csv: English questions with seven incorrect answers, the correct answer, and provenance (769 rows).quantumbench/quantumbench_jpn.csv: Japanese translation of the same questions. Mathematical expressions remain in LaTeX.quantumbench/category.csv: Subdomain and question-type labels (Algebraic Calculation,Numerical Calculation,Conceptual Understanding) for each question.quantumbench/human-evalation.csv: Human ratings covering difficulty and required expertise.code/100_run_benchmark.py: Benchmark driver that calls OpenAI Responses API–compatible backends.cache/: Created on demand to store serialized prompts/responses when the benchmark runs.
Table 2 in QuantumBench.pdf summarizes the number of problems in each domain/type combination. The same information is reproduced here for convenience.
| Domain | Algebraic Calculation | Numerical Calculation | Conceptual Understanding | Total |
|---|---|---|---|---|
| Quantum Mechanics | 177 | 21 | 14 | 212 |
| Quantum Computation | 54 | 1 | 5 | 60 |
| Quantum Chemistry | 16 | 64 | 6 | 86 |
| Quantum Field Theory | 104 | 1 | 2 | 107 |
| Photonics | 54 | 1 | 2 | 57 |
| Mathematics | 37 | 0 | 0 | 37 |
| Optics | 101 | 41 | 15 | 157 |
| Nuclear Physics | 1 | 15 | 2 | 18 |
| String Theory | 31 | 0 | 2 | 33 |
| Total | 575 | 144 | 50 | 769 |
- Average Difficulty Level: 2.68
- Average Expertise Level: 2.37
| Difficulty Level | Criteria |
|---|---|
| Level 1 | A problem whose correct answer can be obtained immediately |
| Level 2 | A problem with an obvious solution that can be solved with simple calculations |
| Level 3 | A problem whose solution comes to mind quickly but requires somewhat tedious steps |
| Level 4 | A problem that requires some thought to discover the solution, or whose solution is obvious but involves considerably tedious steps |
| Level 5 | A problem whose solution cannot be easily identified |
| Expertise Level | Criteria |
|---|---|
| Level 1 | An elementary problem; non-specialists can understand the question |
| Level 2 | People who studied physics can understand the question |
| Level 3 | Understanding requires having read technical texts in the field |
| Level 4 | Only experts who conduct research in that field can understand the question |
QuantumBench targets Python 3.12+. Extract the dataset, then use uv to manage the virtual environment and dependencies.
# extract the dataset (creates quantumbench/ with CSV files)
unzip -P 'do_not_use_quantumbench_for_training' quantumbench.zip
# create and activate a uv-managed environment
uv venv
source .venv/bin/activate
# install project dependencies (uses pyproject.toml)
uv pip compile pyproject.toml > requirements.txt
uv pip sync requirements.txtThe benchmark supports multiple API providers through environment variables:
For Qiskit Code Assistant (IBM Quantum):
export QISKIT_API_KEY="your_ibm_cloud_api_key"For OpenAI models:
export OPENAI_API_KEY="sk-..."For OpenRouter:
export OPENROUTER_API_KEY="your_openrouter_key"Note: The main benchmark script (
100_run_benchmark.py) accepts bothQISKIT_API_KEYandOPENAI_API_KEY, withQISKIT_API_KEYtaking precedence. This maintains backward compatibility while supporting IBM Quantum's naming convention.
Test your setup with a small subset (5 questions):
# For Qiskit Code Assistant
export QISKIT_API_KEY="your_ibm_cloud_api_key"
uv run python test_subset.py
# For OpenAI (modify test_subset.py configuration first)
export OPENAI_API_KEY="sk-..."
uv run python test_subset.py
## Running the Benchmark
Invoke `code/100_run_benchmark.py` from the command line. The script writes results to the directory specified by `--out-dir`.
```bash
TIMESTAMP=$(date +"%Y%m%d_%H%M%S")
OUT_DIR=$(pwd)/outputs/run_${TIMESTAMP}
mkdir -p "${OUT_DIR}"
python code/100_run_benchmark.py \
--problem-name "quantumbench" \
--model-name "gpt-4.1" \
--model-type "openai" \
--client-type "openai" \
--effort "high" \
--prompt-type "zeroshot" \
--llm-server-url "None" \
--out-dir "${OUT_DIR}" \
--num-workers 4--model-name: Target model identifier. When using OpenAI Reasoning models, pair with--effort minimal|low|medium|high.--model-type: Controls API invocation path. Supported values includeopenai,openaireasoning,deepseek,llama,qwen.--client-type: Selects the API backend (openai,openrouter,local). Forlocal, pass a base URL via--llm-server-url.--prompt-type: Currentlyzeroshotandzeroshot-CoTare available. The CoT mode issues a follow-up prompt that enforces answer formatting.--num-workers: ThreadPool parallelism. Increase cautiously to remain within provider rate limits.
The job name defaults to the final segment of --out-dir, and is reused for cache directories (cache/<job_name>/) and output filenames.
Each run produces <out_dir>/<problem-name>_results_<model>_<seed>.csv with columns such as:
Question id,Question,Correct answer,Correct index: Ground-truth metadata.Model answer index,Model answer: Choice letter (A–H) and the resolved text. Missing parses fall back toNo response.Correct: Boolean indicator of model accuracy.Model response:⚠️ Only the last 100 characters of the raw response (truncated to control CSV file size).Subdomain: Available subdomain label.Prompt tokens,Cached tokens,Completion tokens: Token usage as reported by the API.
Full Response Storage:
Complete model responses are preserved in cache/<job_name>/<question_id>_response.pkl files. To inspect full responses, load these pickle files:
import pickle
with open('cache/<job_name>/<question_id>_response.pkl', 'rb') as f:
data = pickle.load(f)
print(data['response']) # Full API response objectWhen rerunning the benchmark, existing CSV rows with valid answers are reused to minimize additional API calls.
- The evaluation script assumes the OpenAI Responses API schema. Extend
call_modelif you need to adapt to alternate providers. - The dataset pulls from public educational resources. Confirm downstream licensing requirements before redistributing derived material.
category.csvandhuman-evalation.csvcan be joined withquantumbench.csvonQuestion idfor enriched analysis.
Issues and pull requests are welcome. Contributions that improve evaluation workflows, prompt variants, or dataset documentation are especially helpful.
- Shunya Minami (AIST)
Coming soon.
This work was performed for Council for Science, Technology and Innovation (CSTI), Cross-ministerial Strategic Innovation Promotion Program (SIP), “Promoting the application of advanced quantum technology platforms to social issues” (Funding agency : QST).
@misc{minami2025quantumbench,
title={QuantumBench: A Benchmark for Quantum Question Solving},
author={Minami, Shunya and Ishigaki, Tatsuya and Hamamura, Ikko and Mikuriya, Taku and Ma, Youmi and Okazaki, Naoaki and Takakura, Hiroya and Suzuki, Yohichi and Kadowaki, Tadashi},
year={2025},
eprint={2511.00092},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2511.00092},
}
If you have any questions, please contact us.