Skip to content

ICLR 2026 JanusVLN

Heungwoo edited this page Jun 1, 2026 · 1 revision

JanusVLN β€” dual implicit memory for Vision-Language Navigation

Venue: ICLR 2026 Β· Authors: Shuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong, Shichao Xie, Xiaolong Wu, Shiyi Liang, Mu Xu, Xing Wei, Ning Guo Β· arXiv: 2509.22548 Β· Category: Embodied navigation / VLN Β· Trend tag: Implicit neural memory for spatial-semantic VLN.

Approach diagram

flowchart LR
  Video[Egocentric RGB video] --> Sem
  Video --> Spa
  Inst[Language instruction] --> LLM
  subgraph Sem["Semantic memory (left brain)"]
    SE[Visual-semantic encoder<br/>retained KV cache]
  end
  subgraph Spa["Spatial memory (right brain)"]
    GE[Spatial-geometric encoder<br/>3D priors from RGB only]
  end
  SE --> Win[Fixed-size sliding-window<br/>implicit memory]
  GE --> Win
  Win --> LLM[Multimodal LLM backbone] --> Act[Navigation action]
Loading

Problem

VLN agents must remember where they have been. Prior methods rely on explicit semantic memory β€” textual cognitive maps or stacks of stored historical frames β€” which causes spatial-information loss, computational redundancy, and unbounded memory growth over long episodes. They also tend to be 2D-semantics-dominant and weak at the 3D spatial reasoning navigation actually requires.

Method

JanusVLN is described as the first VLN framework with dual implicit memory, inspired by human hemispheric specialization (semantic "left brain" + spatial "right brain"):

  • Semantic memory: a visual-semantic encoder maintained as retained key-value caches.
  • Spatial memory: a spatial-geometric encoder that supplies 3D priors β€” crucially, derived from RGB input only, with no supplementary depth or 3D sensor data.

Both memories are kept as compact, fixed-size neural representations and updated with a sliding-window mechanism that reuses prior computation and eliminates redundancy, giving bounded memory cost regardless of episode length. The two streams feed a multimodal LLM backbone that produces navigation actions.

Results

The paper reports relative gains over baselines in the ranges of roughly 10.5–35.5 and 3.6–10.8 on its navigation metrics (per-benchmark SR / SPL / OSR / NE values omitted here pending confirmation from the camera-ready tables). JanusVLN attains its spatial reasoning using only RGB, without explicit 3D inputs.

Significance

JanusVLN argues for steering VLN from 2D-semantics-dominant pipelines toward 3D spatial-semantic synergy using implicit, fixed-size neural memory rather than explicit maps. The dual-memory design is a concrete alternative to map-building for long-horizon embodied agents that need bounded compute and memory.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally