Overview of E-Commerce Bench.
An agent starts with ¥100,000 and runs up to four stores for 365 simulated days. The primary score is the asset multiplier, end-of-year total assets over the opening balance, averaged over 5 independent episodes per model.
<style> .lb-table { border-collapse: collapse; font-size: 13px; width: 100%; margin: 16px 0; } .lb-table th, .lb-table td { padding: 5px 8px; text-align: right; border-bottom: 1px solid #e0e0e0; } .lb-table th { background: #f5f5f5; font-weight: 600; text-align: center; border-bottom: 2px solid #ccc; } .lb-table th:first-child, .lb-table td:first-child { text-align: left; } .lb-table .tier-row td { background: #fafafa; font-style: italic; font-weight: 600; text-align: left; border-bottom: 1px solid #ccc; } .lb-table .best { background: #e8f5e9; font-weight: 700; } .lb-table tr:last-child td { border-bottom: 2px solid #ccc; } </style>| Model | Final assets ¥k, mean | std ¥k | CSE⁺ ↑ | BadSpend% ↓ | Drawdown / peak ↓ | ¥ per tool call ↑ | Controllable return, pp ↓ | AnchorRatio ↓ | Tool calls | Turns | Bankrupt runs |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Proprietary | |||||||||||
| GPT-5.6 Sol (max) | 1,431 | 314 | 0.672 | 18.48 | 0.278 | 363 | 4.46 | 1.217 | 3,668 | 1,367 | 0/5 |
| Fable5 (max) | 805 | 188 | 0.772 | 3.46 | 0.141 | 479 | 0.22 | 1.573 | 1,469 | 945 | 0/5 |
| GPT-5.5 | 702 | 616† | 0.700 | 16.59 | 0.591 | 192 | 0.90 | 1.043‡ | 3,143 | 1,306 | 2/5 |
| Claude Opus 4.8 (max) | 498 | 231 | 0.662 | 5.41 | 0.130 | 266 | 0.20 | 1.309 | 1,497 | 812 | 0/5 |
| Claude Opus 4.7 (max) | 259 | 111 | 0.811 | 0.12 | 0.189 | 156 | 1.75 | 1.421 | 1,023 | 432 | 0/5 |
| Claude Opus 4.6 (max) | 258 | 266 | 0.649 | 14.98 | 0.587 | 129 | 0.76 | 1.268 | 1,221 | 788 | 2/5 |
| Gemini 3.5 Flash | 190 | 79 | 0.777 | 2.13 | 0.450 | 36 | 0.38 | 0.918 | 2,508 | 2,628 | 0/5 |
| Gemini 3.1 Pro | 130 | 130 | 0.660 | 12.58 | 0.688 | 19 | 2.03 | 1.376 | 1,605 | 1,141 | 2/5 |
| Open-weight | |||||||||||
| Qwen3.8-Max-Preview | 416 | 111 | 0.713 | 6.13 | 0.242 | 173 | 0.38 | 0.834 | 1,826 | 962 | 0/5 |
| GLM 5.2 (high) | 301 | 124 | 0.693 | 1.99 | 0.127 | 137 | 0.59 | 1.434 | 1,467 | 672 | 0/5 |
| Kimi K3 | 265 | 110 | 0.632 | 5.87 | 0.247 | 123 | 2.01 | 1.507 | 1,334 | 878 | 0/5 |
| GLM 5.1 | 226 | 192 | 0.662 | 6.49 | 0.231 | 94 | 0.09 | 1.557 | 1,333 | 852 | 0/5 |
| DeepSeek-V4-Pro-Preview (max) | 190 | 100 | 0.654 | 8.99 | 0.166 | 68 | 1.54 | 1.235 | 1,337 | 678 | 0/5 |
| Qwen3.7-Max | 165 | 134 | 0.616 | 9.85 | 0.433 | 59 | −0.05 | 1.433 | 1,108 | 782 | 0/5 |
| GLM 5.2 (max) | 115 | 71 | 0.632 | 19.68 | 0.360 | 10 | 2.24 | 1.362 | 1,430 | 931 | 0/5 |
| Kimi K2.6 | 70 | 14 | 0.596 | 17.22 | 0.377 | −27 | 0.42 | 1.548 | 1,087 | 1,017 | 0/5 |
| Qwen3.6-Plus | 47 | 28 | 0.625 | 10.68 | 0.583 | −40 | 3.03 | 1.333 | 1,325 | 1,023 | 0/5 |
| Qwen3.5-Plus | 1.1 | 11 | 0.625 | 20.11 | 0.979 | −92 | 5.88 | 1.860 | 1,076 | 800 | 4/5 |
Metrics: CSE⁺ = share of the bargaining range captured on honest deals. BadSpend% = procurement cash reaching fraudulent suppliers. AnchorRatio = what a repeat order pays above the agent's own best price with that supplier, over what a random ordering of the same prices would have cost; below 1 means the agent's ordering helped it.
Capability profiles for six of the 18 models. The primary score and the six dimensions run clockwise from the top: profit (mean end-of-year total assets), negotiation (CSE⁺), fraud avoidance (BadSpend%), solvency (drawdown over peak total assets), efficiency (profit per tool call), execution (controllable return rate) and learning (AnchorRatio). Fraud avoidance, solvency, execution and learning are sign-flipped so that higher is better on every axis. Each axis is min-max normalized over the 18 model means, with whiskers over the five episodes and a dashed polygon at the median. No profile fills the polygon.
The rankings diverge across dimensions. The model that ends the year with the most assets captures a smaller share of each bargaining range than four models below it, and routes 18.5% of its procurement spend to fraudulent suppliers against 0.12% for the most cautious. Final assets alone therefore say little about how an agent got there, which is why the suite reports the axes separately.
E-Commerce Bench is a 365-day continuing task. The agent plays a merchant, "Wang Wang", opening up to four online stores, sourcing inventory by negotiating with suppliers, pricing and stocking products, fulfilling orders and handling returns, with one objective: maximize end-of-year assets.
What makes the horizon bite is that nothing resets. Cash spent on inventory is gone until customers pay and escrow settles nine days later; a supplier that overcharged in March is the same supplier in November; and the context window overflows long before day 365, so the agent has to decide what is worth remembering.
Both sides of the market are deterministic, so an outcome difference is attributable to the agent rather than to the environment:
- Demand follows a fixed multi-factor model over data desensitized from a real e-commerce platform: 6,886 products, 60 categories, 12 store types, a year-long calendar of promotions and market shocks.
- Suppliers decide every price through a deterministic negotiation kernel, seeded per (supplier, SKU, cycle). An LLM only renders that decision into dialogue and is not permitted to change it, so no amount of eloquence talks a supplier below its floor. Of 576 suppliers, 152 are fraudulent and run one of five scam patterns, undetectable from price alone by construction.
Install dependencies
pip install -r requirements.txtSet a key for the provider whose model you want to run, plus
OPENAI_API_KEY for the supplier NPC (see below).
export OPENAI_API_KEY=sk-... # provider: openai, and the NPC
export ANTHROPIC_API_KEY=sk-ant-... # provider: anthropic
export GEMINI_API_KEY=... # provider: google
export DASHSCOPE_API_KEY=... # Qwen
export ZHIPU_API_KEY=... # GLM
export MOONSHOT_API_KEY=... # Kimi
export DEEPSEEK_API_KEY=... # DeepSeekmodels_config.json holds the 18 models of the leaderboard above, each at the
reasoning effort it was evaluated with. Pass an entry key to --model:
openai gpt-5.6-sol · gpt-5.5
anthropic claude-fable-5 · claude-opus-4-8 · claude-opus-4-7 · claude-opus-4-6
google gemini-3.5-flash · gemini-3.1-pro
dashscope qwen3.8-max-preview · qwen3.7-max · qwen3.6-plus · qwen3.5-plus
zhipu glm-5.2-high · glm-5.2-max · glm-5.1
moonshot kimi-k3 · kimi-k2.6
deepseek deepseek-v4-pro
The paper's runs reached these models through an internal gateway; the entries
name the same models at their providers' public endpoints. Requests send no
temperature or top_p, and thinking is enabled wherever the family supports it.
To add a model of your own, see
docs/model_providers.md.
npc_tools is a second, separate model — the supplier's role-play voice, by
default gpt-4o-mini. It only renders dialogue into natural language: every
price and accept/reject decision comes from the deterministic kernel, which the
renderer cannot override, so it does not affect the economics. Point it at any
cheap model you have a key for.
python run.py --model gemini-3.5-flash --max-days 10 --max-turns 50 # smoke test, minutes
python run.py --model gemini-3.5-flash # full 365-day episode
python run.py --model gemini-3.5-flash --runs 5 # 5 parallel episodes
bash run.sh # wrapper: live plots, timestamped log dirUseful flags: --max-days, --max-turns, --runs, --initial-balance,
--max-token-capacity, --log-dir. Each run writes to
log/<timestamp>_<model>/ with per-day balances, negotiation metrics, and full
message transcripts.
Analysis of a finished run:
python evaluation/plot_daily_balance.py log/<session>/run_*_daily_balance.csv --output-dir log/<session>/
python evaluation/extract_chatbox.py log/<session>/run_0_messages.jsonl # per-supplier dialogueagent/ turn-based agent loop, LLM client, provider presets, prompts
context_manager/ token-counted context editing for episodes that overflow
tools/ the 18 tools the agent acts through
opponent/ negotiation: kernel, per-(supplier,SKU) instances, fraud, metrics
data/ products, suppliers, categories, events, promotions
evaluation/ plotting and log-analysis scripts
docs/ model_providers.md — how to configure a model
Released under the Apache License 2.0. See LICENSE.
If you find this benchmark useful, please cite:
@misc{fan2026ecommercebenchevaluatingllm,
title = {E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation},
author = {Wei Fan and Xinjie Shen and Xudong Guo and Jianhong Tu and Yang Su and Yinger Zhang and Lianghao Deng and Fengyu Wang and Baohua Dong and Yangqiu Song and Dayiheng Liu},
year = {2026},
eprint = {2608.30730},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2608.30730}
}
