-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 JanusVLN
Venue: ICLR 2026 Β· Authors: Shuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong, Shichao Xie, Xiaolong Wu, Shiyi Liang, Mu Xu, Xing Wei, Ning Guo Β· arXiv: 2509.22548 Β· Category: Embodied navigation / VLN Β· Trend tag: Implicit neural memory for spatial-semantic VLN.
flowchart LR
Video[Egocentric RGB video] --> Sem
Video --> Spa
Inst[Language instruction] --> LLM
subgraph Sem["Semantic memory (left brain)"]
SE[Visual-semantic encoder<br/>retained KV cache]
end
subgraph Spa["Spatial memory (right brain)"]
GE[Spatial-geometric encoder<br/>3D priors from RGB only]
end
SE --> Win[Fixed-size sliding-window<br/>implicit memory]
GE --> Win
Win --> LLM[Multimodal LLM backbone] --> Act[Navigation action]
VLN agents must remember where they have been. Prior methods rely on explicit semantic memory β textual cognitive maps or stacks of stored historical frames β which causes spatial-information loss, computational redundancy, and unbounded memory growth over long episodes. They also tend to be 2D-semantics-dominant and weak at the 3D spatial reasoning navigation actually requires.
JanusVLN is described as the first VLN framework with dual implicit memory, inspired by human hemispheric specialization (semantic "left brain" + spatial "right brain"):
- Semantic memory: a visual-semantic encoder maintained as retained key-value caches.
- Spatial memory: a spatial-geometric encoder that supplies 3D priors β crucially, derived from RGB input only, with no supplementary depth or 3D sensor data.
Both memories are kept as compact, fixed-size neural representations and updated with a sliding-window mechanism that reuses prior computation and eliminates redundancy, giving bounded memory cost regardless of episode length. The two streams feed a multimodal LLM backbone that produces navigation actions.
The paper reports relative gains over baselines in the ranges of roughly 10.5β35.5 and 3.6β10.8 on its navigation metrics (per-benchmark SR / SPL / OSR / NE values omitted here pending confirmation from the camera-ready tables). JanusVLN attains its spatial reasoning using only RGB, without explicit 3D inputs.
JanusVLN argues for steering VLN from 2D-semantics-dominant pipelines toward 3D spatial-semantic synergy using implicit, fixed-size neural memory rather than explicit maps. The dual-memory design is a concrete alternative to map-building for long-horizon embodied agents that need bounded compute and memory.
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model)
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)