-
Notifications
You must be signed in to change notification settings - Fork 0
CVPR 2026 PanoAffordanceNet
Venue: CVPR 2026 (likely β pending CVF virtual-page confirmation) Category: Affordance / Grounding Trend tag: Trend 1 Affiliations: Hunan University (School of AI & Robotics; NERC for Robot Visual Perception & Control) β corresponding author Kailun Yang
flowchart LR
PANO["360Β° panoramic image"] --> DSM["Distortion-Aware<br/>Spectral Modulator"]
DSM --> FEAT["distortion-corrected features"]
FEAT --> OSDH["Omni-Spherical<br/>Densification Head"]
OSDH --> AFFORD["affordance map<br/>full 360Β° coverage"]
Indoor manipulation often happens in scenes where the relevant affordances are spread across 360Β° β a kitchen task may need to grasp something behind, then place it in front. Conventional affordance models work on perspective images and miss the spherical structure.
The setting is weakly supervised affordance grounding under equirectangular projection (ERP), whose core challenges are severe latitude-dependent geometric distortion, semantic dispersion (one affordance spread across multiple disjoint regions), and cross-scale alignment.
- 360-AGD dataset β first high-quality panoramic affordance grounding benchmark; targets 1,919 affordance classes in 360Β° indoor scenes, annotated via keypoint supervision converted to Gaussian heatmaps. Split into an Easy split (512Γ1024 panoramas from 360-Indoor / Gibson) and a Hard split (4552Γ9104 panoramas from PanoContext / Sun360).
- Distortion-Aware Spectral Modulator (DASM) β latitude-dependent spectral calibration that corrects for the panoramic projection's distortion in feature space.
- Omni-Spherical Densification Head (OSDH) β restores topological continuity from sparse activations to produce dense affordance predictions over the full sphere; paired with a Spherical-Aware Hierarchical Decoder.
- Multi-level constraints β pixel-wise localization (BCE), distributional topology consistency (KL divergence), and a regionβtext contrastive (InfoNCE) objective to suppress semantic drift under low supervision.
On the 360-AGD Hard split, PanoAffordanceNet substantially outperforms perspective-image affordance baselines OOAL and OS-AGDO: KLD 1.306 (vs 2.965 / 3.067), SIM 0.474 (vs ~0.10), NSS 4.398 (vs 1.484). On the Easy split it reports KLD 1.270 / SIM 0.506 / NSS 4.490. It also generalizes to perspective views on AGD20K (seen split KLD 0.739 / SIM 0.616 / NSS 1.750).
The panoramic-affordance angle has been underserved β most affordance work is per-view, which fails for whole-scene manipulation planning. PanoAffordanceNet is the cleanest 360Β° affordance baseline at the time of writing.
- arXiv: 2603.09760 (posted 2026-03-10)
- GitHub: GL-ZHU925/PanoAffordanceNet (code + benchmark, "to be released")
- CVPR 2026 survey Β§6
- RealVLG-R1 Β· SaPaVe
β Back to CVPR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)