Skip to content
LegoXPublic

About

Lego-RL: Harness-Native Reinforcement Learning for Coding Agents

Resources

Stars

113 stars

Watchers

1 watching

Forks

Repository files navigation

Lego-RL

Lego-RL

Online reinforcement learning for coding agents in their native harnesses and real repositories.

Paper  ·  Docs  ·  HuggingFace  ·  LegoX  ·  License


Lego-RL is an open-source framework for training coding agents with online reinforcement learning on real software-engineering tasks. It connects Claude Code, OpenHands, and OpenCode to verl while keeping each agent's native control flow.

Each run follows the same loop: an agent works on a repository task in a fresh Harbor sandbox, the task's verifier supplies the reward, and verl updates the policy from the captured trajectory.

News

  • [2026/10/08] Codex training dashboard. The dedicated Codex dashboard shows the cx-6n8g-dlcoenjt6d5urdcu experiment, with training curves and per-trial analysis. Workspace sign-in is required.
  • [2026/09/17] Public training dashboards. We are opening the live dashboards of six real Lego-RL runs, all on Qwen3.5-35B-A3B: GSPO on OpenSWE with each harness natively — OpenHands SDK, OpenCode and Claude Code; SAO (critic-based) on OpenSWE and on a multilingual task set validated on SWE-bench Multilingual; and a mixed-harness run that trains one policy across all three harnesses at once. Every board shows the full reward and validation curves plus per-trial analysis — see Live Training Dashboard.
  • [2026/09/11] Mixed-harness training. A single run can now interleave OpenHands SDK, OpenCode, and Claude Code. Each trajectory resolves one harness from dataset metadata or a weighted policy, while validation stays pinned to one harness so its metrics remain comparable across steps.
  • [2026/09/07] SAO. Single-rollout, critic-based training with decoupled GAE lambdas, a frozen-attention critic, and DIS.
  • [2026/08] First public release. Lego-RL brings Claude Code, OpenHands, and OpenCode into online RL on real repositories, with native harnesses, executable verifier rewards, synchronous or asynchronous training, and live run monitoring.
Lego-RL architecture

Results

Qwen3.5-35B-A3B trained for three epochs (126 steps) on a 2,699-task OpenSWE-derived index, under three native harnesses, and evaluated on the held-out SWE-bench Verified:

Training reward and SWE-bench Verified solve rate across three harnesses

Verifier reward rises under every harness, and every run improves on the benchmark: +6.4 points with OpenHands SDK, +5.8 with Claude Code, +9.4 with OpenCode.

Coding agentModelSWE-bench Verified (%)
OpenHands SDKQwen3.5-35B-A3B64.0
Qwen3.6-35B-A3B67.4
KAT-Coder-V2.5-Dev67.0
Lego-RL-Qwen3.5-35B-A3B70.4 (+6.4)
Claude CodeQwen3.5-35B-A3B62.4
Qwen3.6-35B-A3B63.4
KAT-Coder-V2.5-Dev66.8
Lego-RL-Qwen3.5-35B-A3B68.2 (+5.8)
OpenCodeQwen3.5-35B-A3B57.2
Qwen3.6-35B-A3B60.6
KAT-Coder-V2.5-Dev61.2
Lego-RL-Qwen3.5-35B-A3B66.6 (+9.4)

All numbers are measured by us under the same harness version and evaluation protocol (temperature 0.7, 200 turns, 200k context budget). RL on the 3.5-generation policy beats both the next base generation and KAT-Coder-V2.5-Dev, a model post-trained from it, in all three harnesses. The gains are also harness-specific — KAT-Coder gains 3.4 points under Claude Code, the harness its authors report, but 0.6 under OpenCode and -0.4 under OpenHands SDK — which is exactly why Lego-RL trains inside the harness the agent will actually run in.

Full protocol, ablations, and failure analysis are in the paper.

Live Training Dashboard

Lego-RL training dashboard — per-task solve-rate grid

Seven of our own runs are listed below, all training Qwen3.5-35B-A3B. All seven dashboards are public. The Codex and Mixed harness dashboards open directly without Workspace sign-in:

Run Harness Data (train / val) Algorithm Steps Val (start → best) Dashboard
OpenHands SDK OpenHands SDK OpenSWE 2,699 / SWE-bench Verified 500 GSPO 126 64.0 → 70.4 rl-dashboard-openhands ↗
OpenCode OpenCode OpenSWE 2,699 / SWE-bench Verified 500 GSPO 127 57.2 → 66.6 rl-dashboard-opencode ↗
Claude Code Claude Code OpenSWE 2,699 / SWE-bench Verified 500 GSPO 130 62.4 → 68.2 rl-dashboard-claudecode ↗
Codex Codex OpenSWE 2,699 / SWE-bench Verified 500 GSPO 125 60.0 → 67.6 rl-dashboard-codex ↗
Mixed harness OpenHands SDK + OpenCode + Claude Code OpenSWE 2,699 / SWE-bench Verified 500 GSPO with R3 router replay 126 62.2 → 68.2 rl-dashboard-mixed ↗
SAO on OpenSWE OpenHands SDK OpenSWE 2,699 / SWE-bench Verified 500 SAO 239 64.0 → 68.6 rl-dashboard-sao ↗
SAO multilingual OpenHands SDK Self-made multilingual 1,729 / SWE-bench Multilingual 300 SAO 110 51.7 → 57.0 rl-dashboard-multilingual ↗

Each board follows one run from the first step to the last. The Codex and Mixed harness boards publish read-only training and validation curves; the other boards also provide individual-trial views. Val is the solve rate (%) on the run's validation set, and "best" is the best validation checkpoint. The six original runs use val-core/.../mean@1. Codex uses its registered offline evaluations (200k context, 4,800 seconds per task): 60.0% before training and 67.6% at step 40. Its board displays only cx-6n8g-dlcoenjt6d5urdcu. All runs are fully asynchronous. See the dashboard docs for what each panel shows, or run the dashboard on your own logs:

bash webui/start_dashboard.sh

Key Features

  • Native agents: Claude Code, OpenHands, and OpenCode run through thin adapters. A custom scaffold that speaks the OpenAI or Anthropic API can use the same agent-loop interface.
  • Faithful rollouts: an in-process proxy records token ids, masks, and log-probabilities at generation time. It also handles history rewrites and serves the OpenAI and Anthropic interfaces used by the supported agents.
  • RL and scaling: PPO, GRPO, GSPO, and SAO (single-rollout, critic-based) run on FSDP, VeOmni, or Megatron, either synchronously or fully asynchronously. MoE runs can use R3 routing replay, while trajectory filtering removes broken or over-long rollouts from the loss.
  • Sandboxed rewards: Kubernetes or Docker runs each task in an isolated Harbor environment. The task verifier provides the reward, with image caching and reward-hacking checks available for supported task sets.
  • Data and operations: task indexes, preflight validation, and the optional /rl:check, /rl:run, /rl:status, and /rl:dashboard commands support the complete run lifecycle.
  • Monitoring: the live dashboard combines training curves, validation, per-task results, and individual trajectories. The integration keeps the project glue in src/verl_patch and src/harbor_patch rather than modifying upstream packages.

Getting Started

Tip

Everything from installation and configuration to the full training loop and failure playbook is at lego-rl.pages.dev/docs.

Install and launch the demo run, a real training run at 1/16 scale (8 prompts × 4 responses = 32 trials/step):

git clone https://github.com/LegoX/Lego-RL.git && cd Lego-RL
bash scripts/setup_env.sh                                          # pinned upstreams + self-contained venv

cp scripts/train/examples/demo.env scripts/train/configs/demo.env
$EDITOR scripts/train/configs/demo.env                             # fill the CHANGEME values:
                                                                   #   checkpoint, train/val index, kubeconfig
PREFLIGHT_ONLY=1 bash scripts/train/train.sh scripts/train/configs/demo.env   # validate
bash scripts/train/train.sh scripts/train/configs/demo.env                   # launch

Needs 8× A100/H100-class GPUs, uv, a policy checkpoint, two task indexes, and a reachable Kubernetes cluster (BACKEND=docker drives one machine's daemon instead). The validate step launches nothing. In Claude Code the same run is /rl:run scripts/train/configs/demo.env.

Community

Questions, run reports and contributions are welcome. Scan to join the WeChat group:

License

Apache License 2.0.

Acknowledgement

Built on verl for RL training and Harbor for sandboxed task execution and verifier rewards. Coding agents: Claude Code, OpenHands, and OpenCode.

Citation

@misc{du2026legorlharnessnativereinforcementlearning,
  title={LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents},
  author={Yiming Du and Yuxin Jiang and Tao Yuan and Jianbo Dai and Shaowei Wang and Jierun Chen and Chaofan Tao and Xianzhi Yu and Lifeng Shang and Kam-Fai Wong and Xiaohui Li and Haoli Bai},
  year={2026},
  eprint={2608.17393},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2608.17393},
}

About

Lego-RL: Harness-Native Reinforcement Learning for Coding Agents

Resources

Stars

113 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages