-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 Beyond Binary Success
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Imitation learning 1 Β· paper #76 Authors: David Snyder, Apurva Badithela, Nikolai Matni, George J. Pappas, Anirudha Majumdar, Masha Itkina, Haruki Nishimura arXiv: 2603.13616 Β· program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 frames the problem in three panels: (left) policy comparison arises from design choices β action tokenization, data mixture, architecture, vision encoder β that pit policy Ο_A against Ο_B; (middle) hardware evaluation runs multi-stage trials under a sequential, any-time stopping procedure; (right) N-SCORE delivers comparisons that go beyond binary metrics (partial credit, continuous progress scores), minimize expected trial count, and carry a statistical validity guarantee P[Ο0 > Ο1] β€ Ξ±.
Real-robot evaluation is usually limited to 10β60 rollouts per policy, and comparisons rarely carry statistical guarantees; the state-of-the-art sequential procedure (STEP) is rigorous but only handles binary success, wasting the information in partial-credit rubrics, episodic reward, or trajectory-smoothness metrics that practitioners increasingly use.
The paper introduces N-SCORE, a sequential test built on safe, anytime-valid inference (SAVI). Evidence is aggregated multiplicatively as a test (super)martingale X_{n+1} = (1 + ΞΎ_n(r_{1,n} β r_{0,n}))Β·X_n over paired progress scores, with the null rejected once X exceeds 1/Ξ±*; Ville's inequality gives Type-1 error control at any stopping time (Theorem 1), so evaluators can stop as soon as evidence suffices. The key technical piece is online optimization of the betting fraction ΞΎ_n using kernel-density-estimation-style nonparametric representations of the reward-difference distribution (a family N-SCORE_k indexed by a bandwidth-like parameter k), which handles discrete partial credit and fully nonparametric continuous metrics in one framework β unlike STEP (binary-only) or the parametric ΞΈ-SAVI baseline.
Validation spans over 4,500 hardware rollouts and 2,000 high-fidelity simulation rollouts, including RoboArena/DROID crowd-sourced evaluations and the LBM 1.0 study. On LBM 1.0, partial-credit metrics with sequential testing cut simulation evaluation burden by ~70% (a nearly 1,400-sample reduction versus the 2,000-rollout batch), an over-50% improvement versus sequential binary methods like STEP; on hardware the reduction is ~45% versus batch and 24β30% versus STEP. On synthetic nonparametric data, N-SCOREβ beats the WSR baseline by ~15% in time-to-decision with ~5 points higher power. Fine-grained progress metrics consistently separate competing policies faster than binary success.
Gives the field a statistically rigorous way to certify "policy A beats policy B" with far fewer physical rollouts β and a quantified argument for rubric-based scoring over binary success. Directly relevant to the evaluation-methodology thread in Review-VLA-Evaluation and the LBM co-training comparisons discussed in Review-LBM-Cotraining.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)