Developed and maintained by the Accio team at Alibaba International.
Overview Β· Live leaderboard Β· Mock showcase Β· Quick start Β· Reproducibility Β· Contact
We run models on request, including pre-release and internal builds β and we are open to working together on the benchmark.
RealReplicaBench evaluates whether an agent can complete long-horizon business workflows, not just answer questions about them. Tasks cover browser operations, native-style CLI tools, API/MCP workflows, document and spreadsheet production, public-web research, supplier analysis, product publishing, logistics, and commerce operations. Every task runs in a fresh container and is graded by its own deterministic or LLM-assisted verifier.
- 107 tasks: 53 CLI, 28 browser, 16 file, and 10 API/MCP tasks.
- Three capability slices: 65 text-only, 20 browser-text-capable, and 22 vision-required tasks.
- Stateful evaluation: local mock services model SaaS, commerce, messaging, document, and operational systems without requiring production accounts.
- Auditable outputs: each run preserves the resolved configuration, trajectory, verifier result, artifacts, logs, and container metadata.
The suite uses reproducible local replicas of commerce and business software, so agents must operate interfaces and change state.
![]() |
![]() |
![]() |
| Product publishing Structured catalog and listing operations |
Freight booking Multi-step logistics workflows |
Storefront operations Visual configuration and stateful editing |
Browse 104 rendered pages across eight UI mock services. The showcase is a static visual tour; state-changing interactions run inside the benchmark runtime.
Results are aligned by task_id over the complete 107-task collection. The
tables below are per harness β twelve model families on OpenClaw, thirteen on
Accio β and the twelve present in both are the ones that compare directly.
The published scores were produced through Accio-managed evaluation endpoints
with gemini-3.1-pro-preview as the judge; the public path in this repository
uses bring-your-own credentials.
The live leaderboard is the source of record; the tables below are a snapshot.
Pass and capacity use the same verifier semantics across harnesses. Steps, time, and tokens are descriptive telemetry: tool granularity, runtime scheduling, and provider usage accounting differ, so these values are not normalized efficiency scores.
π₯π₯π₯ mark the top three within each harness. The bar in the Pass column is drawn on a fixed 0β100% scale, not normalized to the leader, so bar lengths are directly comparable between the two tables.
| Model | Pass | Avg. capacity | Avg. steps | Avg. time | Avg. tokens |
|---|---|---|---|---|---|
| π₯ Claude Opus 5 | ββββββββββββββββββββ 60/107 (56.1%) |
0.905 | 47.7 | 12.7 min | 3.47M |
| π₯ Claude Opus 4.8 | ββββββββββββββββββββ 55/107 (51.4%) |
0.860 | 47.6 | 16.4 min | 4.05M |
| π₯ GPT-5.6 Sol | ββββββββββββββββββββ 53/107 (49.5%) |
0.855 | 28.6 | 14.4 min | 2.09M |
| GPT-5.5 | ββββββββββββββββββββ 51/107 (47.7%) |
0.835 | 37.1 | 12.7 min | 2.85M |
| Claude Opus 4.7 | ββββββββββββββββββββ 49/107 (45.8%) |
0.871 | 47.4 | 14.3 min | 4.10M |
| Qwen 3.8 Max Preview | ββββββββββββββββββββ 48/107 (44.9%) |
0.822 | 40.6 | 18.9 min | 2.13M |
| Gemini 3.6 Flash | ββββββββββββββββββββ 48/107 (44.9%) |
0.867 | 46.3 | 13.5 min | 3.28M |
| DeepSeek V4 Flash | ββββββββββββββββββββ 46/107 (43.0%) |
0.827 | 137.8 | 19.2 min | 11.04M |
| GLM 5.2 | ββββββββββββββββββββ 42/107 (39.3%) |
0.814 | 56.9 | 14.8 min | 3.12M |
| Gemini 3.5 Flash | ββββββββββββββββββββ 39/107 (36.4%) |
0.798 | 63.9 | 17.9 min | 5.54M |
| GPT-5.6 Luna | ββββββββββββββββββββ 36/107 (33.6%) |
0.797 | 27.5 | 12.2 min | 1.81M |
| Gemini 3 Flash | ββββββββββββββββββββ 31/107 (29.0%) |
0.744 | 45.1 | 16.1 min | 3.09M |
| Model | Pass | Avg. capacity | Avg. steps | Avg. time | Avg. tokens |
|---|---|---|---|---|---|
| π₯ Claude Opus 5 | ββββββββββββββββββββ 66/107 (61.7%) |
0.861 | 63.2 | 10.1 min | 3.69M |
| π₯ Claude Opus 4.8 | ββββββββββββββββββββ 59/107 (55.1%) |
0.886 | 67.4 | 11.6 min | 4.82M |
| π₯ Claude Opus 4.7 | ββββββββββββββββββββ 56/107 (52.3%) |
0.878 | 61.5 | 6.4 min | 4.32M |
| GPT-5.6 Sol | ββββββββββββββββββββ 55/107 (51.4%) |
0.873 | 53.0 | 5.5 min | 1.85M |
| Qwen 3.8 Max | ββββββββββββββββββββ 52/107 (48.6%) |
0.826 | 67.8 | 15.6 min | 2.93M |
| Gemini 3.6 Flash | ββββββββββββββββββββ 50/107 (46.7%) |
0.815 | 47.7 | 4.6 min | 2.62M |
| GLM 5.2 | ββββββββββββββββββββ 50/107 (46.7%) |
0.787 | 81.0 | 10.8 min | 3.62M |
| DeepSeek V4 Flash | ββββββββββββββββββββ 50/107 (46.7%) |
0.838 | 84.0 | 10.0 min | 5.35M |
| Qwen 3.8 Max Preview | ββββββββββββββββββββ 49/107 (45.8%) |
0.856 | 69.8 | 12.7 min | 2.51M |
| GPT-5.5 | ββββββββββββββββββββ 48/107 (44.9%) |
0.864 | 45.3 | 4.5 min | 1.44M |
| GPT-5.6 Luna | ββββββββββββββββββββ 48/107 (44.9%) |
0.809 | 66.0 | 5.7 min | 2.49M |
| Gemini 3.5 Flash | ββββββββββββββββββββ 46/107 (43.0%) |
0.821 | 91.2 | 9.0 min | 4.80M |
| Gemini 3 Flash | ββββββββββββββββββββ 31/107 (29.0%) |
0.769 | 46.0 | 4.5 min | 2.48M |
The raw task-level result bundles are not stored in Git and do not yet have public immutable URLs or checksums. Until they do, the published board is an audited aggregate keyed by public result IDs, not a standalone reproduction package.
Tip
We run models on request, including pre-release and internal builds, and can evaluate privately against your own checkpoint before you ship it.
We are also open to collaboration β new task domains, mock environments, harness work, or joint evaluation. Tell us what you have in mind.
Prefer to copy rather than click: lianyukun.lyk@alibaba-inc.com Β·
sicong.xsc@alibaba-inc.com
| Metric | Definition |
|---|---|
| Pass | A task passes only when every required verifier check passes; the rate is passes over the 107 aligned tasks. |
| Avg. capacity | Macro mean of each task's checks_passed / checks_total; this preserves partial task completion but is not a weighted official score. |
| Avg. steps | Mean trajectory tool-call count over the displayed task attempts. |
| Avg. time | Mean task wall-clock duration, using summary duration or audited manifest timestamps when the summary duration is zero. |
| Avg. tokens | Mean total model tokens per task after normalizing provider-specific usage fields; cached tokens are included when reported. |
- Docker with Linux container support (
linux/amd64; Apple Silicon hosts can use emulation). - Python 3.11 or newer.
- A model API key and an LLM-judge API key.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
real-replica-bench listThe human-readable tag is mutable, so evaluation commands pin the current release digest:
docker pull --platform linux/amd64 \
acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859The image contains OpenClaw 2026.5.22, the browser stack, and the isolated
domain mock suite.
This example uses Gemini's native generateContent path and the public Google
API:
export GEMINI_API_KEY="..."
real-replica-bench run api-amazon-margin-floor-audit \
--harness openclaw \
--image acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859 \
--platform linux/amd64 \
--openclaw-model google/gemini-3.5-flash \
--openclaw-image-model google/gemini-3.5-flash \
--openclaw-models-config configs/realreplicabench_native_google_direct_models.json \
--llm-judge-provider gemini \
--llm-judge-model gemini-3.1-pro-preview \
--run-id realreplicabench-smokereal-replica-bench run \
--config configs/realreplicabench_openclaw_native_google_direct.yaml \
--run-id "realreplicabench-openclaw-$(date +%Y%m%d-%H%M%S)"Use --limit 1 for a batch-path smoke test. The full suite can be partitioned
with the *_text_only, *_browser_textcapable, and *_vision collection
files under datasets_domain_v1/.
Every route is one config file in configs/, all named
realreplicabench_openclaw<suffix>.yaml. The tables list the suffix.
Managed routes β a provider's own API, billed to that provider's key.
| Route | Suffix | Wire protocol | Credentials |
|---|---|---|---|
| Native Gemini | _native_google_direct |
Gemini generateContent |
GEMINI_API_KEY |
| Native Qwen / DashScope | _qwen37plus_native |
DashScope OpenAI-compatible | DASHSCOPE_API_KEY |
| OpenRouter | (none) | OpenRouter chat, bundled shim | OPENROUTER_API_KEY |
| Qwen through OpenRouter | _qwen37plus_openrouter |
OpenRouter chat, bundled shim | OPENROUTER_API_KEY |
| Custom native Gemini | _native_google |
Gemini generateContent |
Provider-specific |
Bring your own endpoint β point OpenClaw at any base URL you control that speaks one of these four wire formats, and evaluate a self-hosted or pre-release model.
| Wire format | Suffix | Credentials |
|---|---|---|
OpenAI /v1/chat/completions |
_openai_chat |
OPENAI_API_KEY, or your endpoint's var |
OpenAI /v1/responses |
_openai_responses |
OPENAI_API_KEY, or your endpoint's var |
Anthropic /v1/messages |
_anthropic_messages |
ANTHROPIC_API_KEY, or your endpoint's var |
Gemini generateContent |
_custom_gemini |
CUSTOM_GEMINI_BASE_URL + CUSTOM_GEMINI_API_KEY |
Override the endpoint with baseUrl in the models JSON,
--openclaw-provider-base-url (--openclaw-base-url for OpenRouter), or
--openclaw-api to skip the preset entirely β see
docs/openclaw-byo-endpoint.md.
The judge is configured independently of the agent, on Gemini generateContent
or the OpenAI Responses API. Six tasks include LLM-assisted checks; keep the
judge on gemini-3.1-pro-preview unless you report a different one.
Supply credentials through environment variables: the batch runner redacts them
from run.yaml and fails on unresolved ${...} placeholders before a container
starts. Evaluated models run with shell access to their container β see
SECURITY.md for the key-handling rules that implies.
Every route above β the Gemini, Qwen, and OpenRouter agents and both Judges,
including reasoning through the bundled shim and custom upstream base URLs β
has been exercised against local protocol recorders, without real credentials
or billable calls, under the exact request/response contracts covered by
tests/test_public_api.py.
This proves request construction and response parsing, not provider-side model
entitlement, quota, or billing. Before a full run, use --limit 1 with your own
keys and record the provider/model snapshot in the run metadata.
Deeper reference:
docs/openclaw-runtime-image.md for the
runtime image's identity, pin, and customization boundary;
docs/openclaw-native-gemini.md and
docs/openclaw-native-qwen.md for the native
provider routes.
Comparable runs pin these four:
| Component | v1.3.1 pin |
|---|---|
| Task set | realreplicabench_domain_v1_all β 107 task IDs |
| Task definitions | This repository release, including task workspaces and graders |
| Harness | OpenClaw runner in this repository |
| Runtime | acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859 |
Report the rest: provider, exact model and judge identifiers, endpoint class, reasoning configuration, task count, retry policy, and aggregation rule. Compare results only within one benchmark version β a release can change what a task accepts β and never by displayed model name alone: routing, model snapshots, prompt adapters, retry policies, and judge endpoints all change outcomes.
datasets_domain_v1/
βββ realreplicabench_domain_v1_{all,text_only,browser_textcapable,vision}.collection.json
βββ <interface>/<platform>/<task>/
βββ task.toml task.md workspace/ agent-visible
βββ grader/ services/ private/ rubric.json
Only task.md and workspace/ are staged into the agent-visible task tree;
graders, rubrics, private seeds, service launchers, and mock source stay
outside it, and final artifacts go to /task/outputs/. After the agent exits,
the host-side verifier reads those outputs and the isolated mock state, writes
the reward record, archives logs and trajectories, removes the container, and
leaves:
runs/<run_id>/
βββ run.yaml summary.json summary.md report.html
βββ tasks/<index>-<task_id>/
βββ manifest.json agent/ verifier/ workspace/outputs/ screenshots/ container/
We are asking for your mock environments. A benchmark with a fixed task set
decays: models saturate it and its answers drift into training data. Each new
replica service β a real service's API semantics, state transitions, and above
all its rejections, running offline and scored deterministically β is a family
of tasks no model has been trained on. The fourteen shipping today are
registered in real_replica_bench/mock_services/registry.py.
CONTRIBUTING.md states the bar a new mock has to clear, and
the rules for task fixes, graders, and harness changes. A merged mock reaches
the published benchmark when maintainers next rebake the runtime image.
The Accio team at Alibaba International built the harness, the mock services,
and the v1 task suite; your pull request adds you to
CONTRIBUTORS.md alongside the mock itself.
Report vulnerabilities privately per SECURITY.md; third-party
provenance is inventoried in
THIRD_PARTY_NOTICES.md.
Citation metadata is available in CITATION.cff. Cite
RealReplicaBench (Accio) together with release v1.3.1 and the exact Git
commit used for evaluation. Until the accompanying paper is published, cite
the repository directly:
@misc{Lian2026RealReplicaBench,
author={Yukun Lian and Lei Wei and Sicong Xie and Guannan Zhang and Kesu
Wang and Hongyu Li and Chenhao Jiang and Lanbo Lin and Tianyuan
Yang and Xiaoyu Guo and Li Cai and Jialong Zhu},
title={RealReplicaBench: A Stateful Agent Benchmark for Long-Horizon Commerce and Business Workflows},
note={GitHub repository, v1.3.1},
howpublished={\url{https://github.com/Accio-Lab/RealReplicaBench}},
year={2026}
}RealReplicaBench is open source. It ships under two licenses, split the same way as the repository itself:
| Scope | License | File |
|---|---|---|
| Harness, Python package, mock-service code, scripts, and configs | Apache License 2.0 | LICENSE |
Task suite under datasets_domain_v1/ (task definitions, workspaces, graders, rubrics) |
Creative Commons Attribution 4.0 International (CC BY 4.0) | LICENSE-DATA |
Commercial use is allowed. Keep the license and attribution notices, state significant changes, and credit the benchmark as described under Citation. Neither license grants trademark rights β "Accio" and "RealReplicaBench" identify this benchmark, not a fork of it.
Important
These terms cover Accio's own contributions only. The repository also
contains mirrored stylesheets, webfonts, icons, and recorded API responses
whose rights their owners retain. Every one is inventoried by owner and path
in THIRD_PARTY_NOTICES.md; read it before
redistributing.


