Skip to content

ICLR 2026 ManipEvalAgent

Heungwoo edited this page Jun 1, 2026 · 1 revision

ManipEvalAgent β€” promptable, agentic evaluation of manipulation policies

Venue: ICLR 2026 Β· Authors: Yiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang, Xiangyu Zhao, Qingyao Wu Β· Paper: OpenReview https://openreview.net/forum?id=3u6AkbWEls (no arXiv preprint located as of 2026-06) Β· Category: Benchmark / evaluation framework Β· Trend tag: Agentic, VLM-in-the-loop policy evaluation

Approach diagram

flowchart LR
  Q[User query / instruction<br/>"how robust is this policy to clutter?"] --> Plan{Evaluation agent<br/>code-gen + planner}
  Plan --> Batch[Run small batch of<br/>rollouts in sim]
  Batch --> VLM[VLM video understanding<br/>per-rollout analysis]
  VLM --> Obs[Intermediate observations<br/>failure modes, sub-goal hits]
  Obs --> Plan
  Plan -->|stopping criterion met| Report[Fine-grained diagnostic report<br/>not just a single scalar]
Loading

Problem

Standard simulation benchmarks evaluate manipulation policies by large-scale sampling β€” hundreds to thousands of rollouts per task β€” which is slow, and they collapse the outcome into a single scalar success rate that hides why a policy fails. They also use a fixed, non-adaptive pipeline that cannot answer targeted user questions (e.g. robustness to a specific distractor, or instruction-following under paraphrase).

Method

ManipEvalAgent reframes evaluation as an agentic, multi-round process:

  • Promptable. The agent ingests a natural-language user query and plans the evaluation procedure (which tasks, perturbations, and metrics) accordingly, via code generation.
  • Small-batch, adaptive sampling. Instead of exhaustive sampling, it runs small batches of rollouts and adaptively plans the next round based on intermediate observations, concentrating effort where uncertainty or failure is high.
  • VLM video understanding. A vision-language model parses rollout videos to produce user-instruction-centric, fine-grained analysis β€” diagnostic text about failure causes rather than a bare number.

Results

The framework reportedly reaches conclusions comparable to large-scale simulation benchmarks while significantly shortening total evaluation time, and delivers interpretable, diagnostic output beyond a scalar score. (The OpenReview record does not expose specific benchmark names or numeric speedups in the public abstract; numbers omitted here pending the full PDF.)

Significance

Pushes policy evaluation β€” not just policy learning β€” toward the agentic / VLM-in-the-loop paradigm. Complements heavy fixed-protocol simulators by offering a cheap, query-driven, interpretable alternative, in the same spirit as world-model-as-evaluator and real-to-sim benchmarking efforts at this venue.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally