Repository navigation
Simulation RL Environments, part 1: PortSimEnv v1. Ideas for v2, v3, post-training and data #36
adithya-s-k
started this conversation in
Ideas
Replies: 1 comment
This fits into what we support in https://www.agentenvframework.com/ . Would love to collab on this! Happy to do a prototype as a showcase |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Simulation RL Environments is an experimental, work-in-progress series: take real-world work, rebuild it as a deterministic simulation from its own records, and train agents in it with RL.
Part 1 is PortSimEnv v1: re-plan a week of container-ship dockings at a Port of Barcelona quay, built from the port's 2024 records (1,784 real calls), with gales, crane breakdowns, closures, emergencies and diverted traffic. One graded submit per episode, scored deterministically against a CP-SAT-proven optimum.
Eval on 50 held-out weeks (mean reward): GPT-6.1 Sol 0.89 · Claude Sonnet 5.5 0.78 · GLM-5.3-Flash 0.47 · Qwen3.8-2.4T 0.38 · GLM-5.3 0.31 · Qwen3.8-27B 0.21.
The open question: from seed data we can build realistic environments, but how do we make them as close to the real world as possible? Ideas, pointers and critiques are welcome on:
All reactions