Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

arXiv Website Cite Python

Mean end-of-year total assets, eighteen models, five 365-day episodes each

Overview of E-Commerce Bench.

Results

An agent starts with ¥100,000 and runs up to four stores for 365 simulated days. The primary score is the asset multiplier, end-of-year total assets over the opening balance, averaged over 5 independent episodes per model.

<style> .lb-table { border-collapse: collapse; font-size: 13px; width: 100%; margin: 16px 0; } .lb-table th, .lb-table td { padding: 5px 8px; text-align: right; border-bottom: 1px solid #e0e0e0; } .lb-table th { background: #f5f5f5; font-weight: 600; text-align: center; border-bottom: 2px solid #ccc; } .lb-table th:first-child, .lb-table td:first-child { text-align: left; } .lb-table .tier-row td { background: #fafafa; font-style: italic; font-weight: 600; text-align: left; border-bottom: 1px solid #ccc; } .lb-table .best { background: #e8f5e9; font-weight: 700; } .lb-table tr:last-child td { border-bottom: 2px solid #ccc; } </style> Table 1: End-of-year leaderboard, 18 models, five episodes each, sorted by mean final assets within tier. Tinted cells mark the best scored column. †GPT-5.5's spread is a population estimate, ¥689k on the sample estimator. ‡GPT-5.5's AnchorRatio averages three episodes, its two bankrupt runs opening no repeat order with a supplier it had already dealt with.
Model Final assets ¥k, mean std ¥k CSE⁺ ↑ BadSpend% ↓ Drawdown / peak ↓ ¥ per tool call ↑ Controllable return, pp ↓ AnchorRatio ↓ Tool calls Turns Bankrupt runs
Proprietary
GPT-5.6 Sol (max) 1,431 314 0.672 18.48 0.278 363 4.46 1.217 3,668 1,367 0/5
Fable5 (max) 805 188 0.772 3.46 0.141 479 0.22 1.573 1,469 945 0/5
GPT-5.5 702 616† 0.700 16.59 0.591 192 0.90 1.043‡ 3,143 1,306 2/5
Claude Opus 4.8 (max) 498 231 0.662 5.41 0.130 266 0.20 1.309 1,497 812 0/5
Claude Opus 4.7 (max) 259 111 0.811 0.12 0.189 156 1.75 1.421 1,023 432 0/5
Claude Opus 4.6 (max) 258 266 0.649 14.98 0.587 129 0.76 1.268 1,221 788 2/5
Gemini 3.5 Flash 190 79 0.777 2.13 0.450 36 0.38 0.918 2,508 2,628 0/5
Gemini 3.1 Pro 130 130 0.660 12.58 0.688 19 2.03 1.376 1,605 1,141 2/5
Open-weight
Qwen3.8-Max-Preview 416 111 0.713 6.13 0.242 173 0.38 0.834 1,826 962 0/5
GLM 5.2 (high) 301 124 0.693 1.99 0.127 137 0.59 1.434 1,467 672 0/5
Kimi K3 265 110 0.632 5.87 0.247 123 2.01 1.507 1,334 878 0/5
GLM 5.1 226 192 0.662 6.49 0.231 94 0.09 1.557 1,333 852 0/5
DeepSeek-V4-Pro-Preview (max) 190 100 0.654 8.99 0.166 68 1.54 1.235 1,337 678 0/5
Qwen3.7-Max 165 134 0.616 9.85 0.433 59 −0.05 1.433 1,108 782 0/5
GLM 5.2 (max) 115 71 0.632 19.68 0.360 10 2.24 1.362 1,430 931 0/5
Kimi K2.6 70 14 0.596 17.22 0.377 −27 0.42 1.548 1,087 1,017 0/5
Qwen3.6-Plus 47 28 0.625 10.68 0.583 −40 3.03 1.333 1,325 1,023 0/5
Qwen3.5-Plus 1.1 11 0.625 20.11 0.979 −92 5.88 1.860 1,076 800 4/5

Metrics: CSE⁺ = share of the bargaining range captured on honest deals. BadSpend% = procurement cash reaching fraudulent suppliers. AnchorRatio = what a repeat order pays above the agent's own best price with that supplier, over what a random ordering of the same prices would have cost; below 1 means the agent's ordering helped it.

Capability profiles for six of the 18 models

Capability profiles for six of the 18 models. The primary score and the six dimensions run clockwise from the top: profit (mean end-of-year total assets), negotiation (CSE⁺), fraud avoidance (BadSpend%), solvency (drawdown over peak total assets), efficiency (profit per tool call), execution (controllable return rate) and learning (AnchorRatio). Fraud avoidance, solvency, execution and learning are sign-flipped so that higher is better on every axis. Each axis is min-max normalized over the 18 model means, with whiskers over the five episodes and a dashed polygon at the median. No profile fills the polygon.

The rankings diverge across dimensions. The model that ends the year with the most assets captures a smaller share of each bargaining range than four models below it, and routes 18.5% of its procurement spend to fraudulent suppliers against 0.12% for the most cautious. Final assets alone therefore say little about how an agent got there, which is why the suite reports the axes separately.

Overview

E-Commerce Bench is a 365-day continuing task. The agent plays a merchant, "Wang Wang", opening up to four online stores, sourcing inventory by negotiating with suppliers, pricing and stocking products, fulfilling orders and handling returns, with one objective: maximize end-of-year assets.

What makes the horizon bite is that nothing resets. Cash spent on inventory is gone until customers pay and escrow settles nine days later; a supplier that overcharged in March is the same supplier in November; and the context window overflows long before day 365, so the agent has to decide what is worth remembering.

Both sides of the market are deterministic, so an outcome difference is attributable to the agent rather than to the environment:

  • Demand follows a fixed multi-factor model over data desensitized from a real e-commerce platform: 6,886 products, 60 categories, 12 store types, a year-long calendar of promotions and market shocks.
  • Suppliers decide every price through a deterministic negotiation kernel, seeded per (supplier, SKU, cycle). An LLM only renders that decision into dialogue and is not permitted to change it, so no amount of eloquence talks a supplier below its floor. Of 576 suppliers, 152 are fraudulent and run one of five scam patterns, undetectable from price alone by construction.

Environment Setup

Install dependencies

pip install -r requirements.txt

Set a key for the provider whose model you want to run, plus OPENAI_API_KEY for the supplier NPC (see below).

export OPENAI_API_KEY=sk-...        # provider: openai, and the NPC
export ANTHROPIC_API_KEY=sk-ant-... # provider: anthropic
export GEMINI_API_KEY=...           # provider: google
export DASHSCOPE_API_KEY=...        # Qwen
export ZHIPU_API_KEY=...            # GLM
export MOONSHOT_API_KEY=...         # Kimi
export DEEPSEEK_API_KEY=...         # DeepSeek

Experiment Configuration

models_config.json holds the 18 models of the leaderboard above, each at the reasoning effort it was evaluated with. Pass an entry key to --model:

openai      gpt-5.6-sol · gpt-5.5
anthropic   claude-fable-5 · claude-opus-4-8 · claude-opus-4-7 · claude-opus-4-6
google      gemini-3.5-flash · gemini-3.1-pro
dashscope   qwen3.8-max-preview · qwen3.7-max · qwen3.6-plus · qwen3.5-plus
zhipu       glm-5.2-high · glm-5.2-max · glm-5.1
moonshot    kimi-k3 · kimi-k2.6
deepseek    deepseek-v4-pro

The paper's runs reached these models through an internal gateway; the entries name the same models at their providers' public endpoints. Requests send no temperature or top_p, and thinking is enabled wherever the family supports it. To add a model of your own, see docs/model_providers.md.

npc_tools is a second, separate model — the supplier's role-play voice, by default gpt-4o-mini. It only renders dialogue into natural language: every price and accept/reject decision comes from the deterministic kernel, which the renderer cannot override, so it does not affect the economics. Point it at any cheap model you have a key for.

Running Experiments

python run.py --model gemini-3.5-flash --max-days 10 --max-turns 50   # smoke test, minutes
python run.py --model gemini-3.5-flash                               # full 365-day episode
python run.py --model gemini-3.5-flash --runs 5                      # 5 parallel episodes
bash run.sh                                                          # wrapper: live plots, timestamped log dir

Useful flags: --max-days, --max-turns, --runs, --initial-balance, --max-token-capacity, --log-dir. Each run writes to log/<timestamp>_<model>/ with per-day balances, negotiation metrics, and full message transcripts.

Analysis of a finished run:

python evaluation/plot_daily_balance.py log/<session>/run_*_daily_balance.csv --output-dir log/<session>/
python evaluation/extract_chatbox.py log/<session>/run_0_messages.jsonl   # per-supplier dialogue

Repository Structure

agent/            turn-based agent loop, LLM client, provider presets, prompts
context_manager/  token-counted context editing for episodes that overflow
tools/            the 18 tools the agent acts through
  opponent/       negotiation: kernel, per-(supplier,SKU) instances, fraud, metrics
data/             products, suppliers, categories, events, promotions
evaluation/       plotting and log-analysis scripts
docs/             model_providers.md — how to configure a model

License

Released under the Apache License 2.0. See LICENSE.

Cite

If you find this benchmark useful, please cite:

@misc{fan2026ecommercebenchevaluatingllm,
  title         = {E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation},
  author        = {Wei Fan and Xinjie Shen and Xudong Guo and Jianhong Tu and Yang Su and Yinger Zhang and Lianghao Deng and Fengyu Wang and Baohua Dong and Yangqiu Song and Dayiheng Liu},
  year          = {2026},
  eprint        = {2608.30730},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2608.30730}
}

About

Long-horizon benchmark where 18 LLM agents got ¥100,000 each and ran simulated online stores for 365 days on real market data: negotiating with suppliers, pricing, managing inventory, keeping cash flow alive.

Resources

Stars

66 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages