Skip to content

SO-101 Action-Space Benchmark v1.0.0

Latest

Choose a tag to compare

@matthewwoodc0 matthewwoodc0 released this 08 Aug 18:58
bc2e032

SO-101 Action-Space Benchmark v1.0.0 is the public release of the completed EFF-001 study.

Result

Across five demonstration budgets, three nested dataset ladders, and five model seeds, joint-delta policies achieved normalized success AUC 0.624306. End-effector tool-delta policies achieved 0.370139. The paired difference was 0.254167 with a crossed-factor 95% interval of [0.148958, 0.353125]. All 15 paired AUC units favored joint-delta.

Scope

This result applies only to the registered MuJoCo pickup design. It does not show that joint actions are universally better. It does not measure hardware behavior or evaluate a vision policy, language conditioning, or a VLA. Locked evaluation and protocol-v2 final data were not accessed.

Reproduce

Run make verify-result to check the compact public evidence, recompute the endpoint and interval, and rebuild the tables and figures. Run make smoke for a small closed-loop training and rollout check.

Evidence

The complete 946-file raw evidence archive is public on Zenodo: https://doi.org/10.5281/zenodo.21854125

Archive SHA-256: 24e6b1e7e28ba6bfa97d8bed11668d3730bc9655a1bdc04607e1c8a587f735f1

See CHANGELOG.md for the full release contents.