Skip to content

ICLR 2026 REI Bench

Heungwoo edited this page Jun 1, 2026 · 2 revisions

REI-Bench β€” Vague Referring Expressions in Embodied Task Planning

Full title: REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning? Venue: ICLR 2026 Β· OpenReview: vmBIF25KLf Β· arXiv: 2505.10872 Authors: Chenxi Jiang, Chuhao Zhou, Jianfei Yang (MARS Lab, School of Mechanical & Aerospace Engineering, Nanyang Technological University) Category: Benchmark β€” Instruction following / task planning Trend tag: Evaluation / pragmatics for embodied LLM planners

Approach diagram

flowchart LR
  User[Vague human instruction<br/>referring expression] --> Plan[LLM task planner]
  Plan --> Fail[Failure mode:<br/>missing objects in plan]
  User --> TOCC[Task-oriented<br/>context cognition]
  TOCC --> Clear[Clarified instruction]
  Clear --> Plan
Loading

Problem

LLM-based robot task planners are typically evaluated on clean, well-specified instructions, but real users β€” especially non-experts, the elderly, and children β€” produce vague referring expressions ("that thing over there", "the one I used yesterday"). No benchmark systematically measures how planners degrade under this vagueness.

Method

REI-Bench is the first benchmark that systematically models vague referring expressions grounded in pragmatic theory for embodied task planning. It is built on ALFRED, selecting six household tasks (Pick & Place, Stack & Place, Clean & Place, Heat & Place, Cool & Place, Examine in Light), filtered to instances solvable with clear instructions by LLaMA3.1-8B + SayCan.

Instructions are organized into three levels of referential difficulty by the ratio of explicit to implicit REs:

  • Explicit REs β€” original expressions preserved (e.g., "potato").
  • Mixed REs β€” task instruction uses implicit REs while the dialogue context still carries explicit ones.
  • Implicit REs β€” all expressions replaced by pronouns/descriptors (e.g., "the heated one").

Evaluation spans LLMs (GPT-4o-mini, LLaMA3.1-8B, Ministral-8B, Gemma2-9B, DeepSeek-Math-7B, Qwen2.5-7B) across planning frameworks (SayCan, DAG-Plan, HPE, LLM+P).

The authors also propose task-oriented context cognition, a prompting/reasoning approach that rewrites vague user input into clearer instructions before planning.

Results

Vagueness in referring expressions causes success-rate drops of up to 36.9% for current LLM planners; analysis attributes most failures to missing objects in the generated plan. Task-oriented context cognition achieves SOTA versus aware-prompting, chain-of-thought, and in-context learning baselines (specific final-success numbers not stated in abstract).

Significance

Reframes embodied-LLM evaluation around pragmatic robustness rather than literal-instruction following β€” a prerequisite for deployment to non-expert users. Sits alongside RoboArena in the 2026 push toward more realistic embodied evaluation.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally