Overview of WebDevJudge, including data collection, annotation, and evaluation.
- 2026-02-11: We are currently working on refining some ambiguous labels to ensure higher consistency (done) and scaling up the benchmark to over 1k instances. The updated version is scheduled for release in early March 2026. All inference results of the models are available here.
- 2025-11-16: We release the results for WebDevJudge Unit (UI-TARS 1.5) using code in this repository. You can find the results and corresponding data link in webdevjudge_unit/README.md.
- 2025-11-10: We release the full WebDevJudge Unit dataset.
- 2025-10-21: We release our paper and data. Check it out!
WebDevJudge is a benchmark for evaluating the performance of LLM-as-a-judge for web development tasks. It supports both static and interactive assessment of web development quality with high-quality preference labels.
First, install the required Python dependencies:
pip install -r requirements.txtNext, populate the api_keys/config.json file with your API keys, following this format:
{
# <model_name> is the name you specify in the command line; other fields are for the OpenAI SDK.
"model_name": {
"model": "model_name",
"base_url": "base_url",
"api_key": "api_key"
}
}THIS REPOSITORY DOES NOT CONTAIN ANY RAW DATA FILES. We only provide the subset labeled with index (data/index2label.json). The indices provided in this repository are for citation and evaluation purposes only and do not constitute a disclosure or transfer of the original dataset.
WebDevJudge is built upon the webdev-arena-preference-10k dataset. Due to strict licensing restrictions, you must independently acquire the original dataset and comply with all terms set forth by the original data provider.
By using this benchmark, you acknowledge and agree to the following terms:
- Independent Acquisition: You are solely responsible for obtaining the original
webdev-arena-preference-10kdataset from its official source. - Full Compliance: You must adhere to all terms and conditions of the original dataset's license.
We provide a script to download the original dataset and prepare it for the benchmark. Run the following command:
python data/prepare.pyAfter execution, the data and category information will be saved to data/all.jsonl and data/category.csv, respectively.
Setup instructions for the interactive environment are available in envs/README.md.
Currently supported platforms:
- CentOS
Support for additional environments (e.g., Ubuntu, Docker) is forthcoming.
To perform a static code evaluation, run the following command:
bash run.shLLM responses will be saved in JSON Lines format to the outputs directory, and predictions will be saved in CSV format to the results directory.
Alternatively, you can evaluate existing results directly with this command:
python run.py \
--setting likert \
--mode pair \
--model gpt-4.1The parameters for run.py are explained in detail below:
--setting: The evaluation setting. Options:likert(default),rubric.--mode: The evaluation mode. Options:pair(default),single.--with_image: Include screenshots in the evaluation. Screenshots must be generated beforehand using the script described in check/README.md.--eval: Evaluate existing results without running a new evaluation.--data_path: Path to the dataset.--screenshots_dir: Path to the directory containing screenshots.--output_dir: Path to the directory where results will be saved.--model: The model to use for evaluation, which must correspond to a key inapi_keys/config.json.--rubric_type: The type of rubric to use. Options:combined(default),static,dynamic,intention.--rubric_path: Path to the rubric file.
Note that to use the rubric-based evaluation protocol, you must first generate the rubric. Run the following command:
python data/process_rubric.py --model gpt-4.1The generated rubric will be saved to data/rubric.jsonl, enabling you to run evaluations with the rubric setting.
Before proceeding, ensure the interactive environment is set up and the rubric has been generated. Then, follow these steps to run the dynamic interactive evaluation:
-
Create a Next.js environment for website deployment:
cd WebDevJudge bash envs/set_up_nextjs_env.sh workspace/workspace_<worker_id>
Here,
<worker_id>is a unique, zero-indexed identifier for each worker. For parallel evaluation, create a separate environment for each worker with sequential IDs (e.g., 0, 1, 2). During the Next.js setup, select the default options when prompted. Ensure theworkspacedirectory is at the same level as therun_serial.shandrun_parallel.shscripts; otherwise, update the paths within the scripts accordingly. -
Process the rubric and save intermediate results:
python agent.py --do_process --base_dir /data/WebDevJudge_test --add_rubric
Intermediate results will be saved to
/data/WebDevJudge_test, and website paths will be written towebs.txt. Each line in this file corresponds to a website path. The directory structure withinbase_diris as follows:<question_id>/<model_id>/ index.tsx intention/ part1/ metadata.json ... static/ part1/ metadata.json ... dynamic/ basic/ part1/ metadata.json ... complex/ part1/ metadata.json ... ... tasks.txt metadata.jsonIf you want to deploy the website manually. You can copy the
index.tsxfile to theworkspace/workspace_<worker_id>/pages/index.tsxdirectory and run the following command to start the Next.js server.cd workspace/workspace_<worker_id> npm run dev -- -p 3000
-
Run the dynamic interactive evaluation:
For a single worker, run the following command:
bash run_serial.sh <display_port> <start_line> <end_line> <worker_id>
<display_port>: The Xvfb display port.<start_line>,<end_line>: The start and end lines inwebs.txtto process (1-indexed).<worker_id>: The worker ID.
Example:
bash run_serial.sh 99 1 10 0
All logs will be printed directly to the console. To validate your environment setup, you can run an evaluation on a single website:
bash run_serial.sh 99 921 921 0
To run the evaluation in parallel across multiple workers, use the following command:
bash run_parallel.sh <num_chunks> <start_display> <dataset_start> <dataset_end>
<num_chunks>: The number of workers to use.<start_display>: The starting display port (e.g., the port forworkspace_0).<dataset_start>,<dataset_end>: The range of lines inwebs.txtto process (1-indexed).
Example:
bash run_parallel.sh 4 99 1 100
Logs for parallel runs will be saved in the
parallel_logsdirectory. -
Evaluate the results:
Once the evaluation is complete, run the following command to process the results, before running the following command, please make sure the
webs.txtfile exists:python agent.py --base_dir /data/WebDevJudge_test --do_eval
The final predictions will be saved in CSV format to the
resultsdirectory. If you want to evaluate a subset of the results, you can modify thewebs.txtfile to pick the websites and run the evaluation again.
Webdevjudge Unit is a task-level benchmark for assessing the capability of evaluators to verify task feasibility. Each instance in the benchmark contains a html code, a task instruction, an expected result, and a label indicating whether the task is feasible.
Detailed instruction to run the WebDevJudge Unit is available in webdevjudge_unit/README.md.
If you find this work useful, please consider citing our work. We truly appreciate your support!
@misc{li2026webdevjudgeevaluatingmllmscritiques,
title={WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality},
author={Chunyang Li and Yilun Zheng and Xinting Huang and Tianqing Fang and Jiahao Xu and Lihui Chen and Yangqiu Song and Han Hu},
year={2026},
eprint={2510.18560},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2510.18560},
}Chunyang Li (cliei@connect.ust.hk)
