-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 CompassNav
Venue: ICLR 2026 Β· Authors: LinFeng Li, Jian Zhao, Yuan Xie, Xin Tan, Xuelong Li Β· arXiv: 2510.10154 Β· Category: Embodied navigation / VLN Β· Trend tag: Decision-understanding training (SFT+RFT) for navigation LVLMs.
flowchart LR
Expert[Single GT path<br/>R2R-style data] --> Astar[A* geodesic distances<br/>annotate ALL feasible actions]
Astar --> Data[Compass-Data-22k<br/>dense field of correctness]
Data --> SFT[Stage 1: SFT<br/>navigational priors]
SFT --> RFT[Stage 2: RFT<br/>Gap-Aware Hybrid Reward]
RFT --> LVLM[7B open-source LVLM<br/>weighs options, then decides]
LVLM --> Act[Navigation action]
Standard navigation training reduces the task to sequence-to-sequence replication of a single correct path. Datasets like R2R provide only one ground-truth trajectory, forcing rigid path imitation that lacks the counterfactual data needed for robust decision-making β the model memorizes routes instead of learning to weigh options and decide.
CompassNav distills spatial reasoning and decision-making directly into open-source LVLMs:
- Compass-Data-22k: a 22k-trajectory dataset whose RFT subset annotates all feasible actions at each step using A* geodesic distances, producing a dense field of correctness across the decision space rather than a single labeled move.
- Gap-Aware Hybrid Reward: a dynamic reward that adapts to decision certainty β decisive signals for clearly optimal actions, nuanced scores to encourage exploration when options are close.
- Two-stage recipe: preparatory SFT to instill navigational priors, then Reward-based Fine-Tuning (RFT) so the agent learns relative move quality instead of route memorization.
A 7B CompassNav model reaches state-of-the-art on goal-navigation benchmarks, outperforming much larger general-purpose models including GPT-4o and surpassing o1-mini, and the paper demonstrates real-world robot navigation. (Exact SR / SPL figures omitted here pending the camera-ready tables.)
CompassNav reframes navigation training as decision understanding rather than path imitation: by densely supervising the full action space with geodesic correctness and a certainty-aware reward, a compact open LVLM can beat frontier proprietary models on embodied goal navigation.
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model)
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)