Skip to content

CVPR 2026 PanoAffordanceNet

Heungwoo edited this page Jun 1, 2026 · 2 revisions

PanoAffordanceNet β€” Holistic Affordance Grounding in 360Β° Indoor Environments

Venue: CVPR 2026 (likely β€” pending CVF virtual-page confirmation) Category: Affordance / Grounding Trend tag: Trend 1 Affiliations: Hunan University (School of AI & Robotics; NERC for Robot Visual Perception & Control) β€” corresponding author Kailun Yang

Approach diagram

flowchart LR
  PANO["360Β° panoramic image"] --> DSM["Distortion-Aware<br/>Spectral Modulator"]
  DSM --> FEAT["distortion-corrected features"]
  FEAT --> OSDH["Omni-Spherical<br/>Densification Head"]
  OSDH --> AFFORD["affordance map<br/>full 360Β° coverage"]
Loading

Problem

Indoor manipulation often happens in scenes where the relevant affordances are spread across 360Β° β€” a kitchen task may need to grasp something behind, then place it in front. Conventional affordance models work on perspective images and miss the spherical structure.

The setting is weakly supervised affordance grounding under equirectangular projection (ERP), whose core challenges are severe latitude-dependent geometric distortion, semantic dispersion (one affordance spread across multiple disjoint regions), and cross-scale alignment.

  • 360-AGD dataset β€” first high-quality panoramic affordance grounding benchmark; targets 1,919 affordance classes in 360Β° indoor scenes, annotated via keypoint supervision converted to Gaussian heatmaps. Split into an Easy split (512Γ—1024 panoramas from 360-Indoor / Gibson) and a Hard split (4552Γ—9104 panoramas from PanoContext / Sun360).
  • Distortion-Aware Spectral Modulator (DASM) β€” latitude-dependent spectral calibration that corrects for the panoramic projection's distortion in feature space.
  • Omni-Spherical Densification Head (OSDH) β€” restores topological continuity from sparse activations to produce dense affordance predictions over the full sphere; paired with a Spherical-Aware Hierarchical Decoder.
  • Multi-level constraints β€” pixel-wise localization (BCE), distributional topology consistency (KL divergence), and a region–text contrastive (InfoNCE) objective to suppress semantic drift under low supervision.

Results

On the 360-AGD Hard split, PanoAffordanceNet substantially outperforms perspective-image affordance baselines OOAL and OS-AGDO: KLD 1.306 (vs 2.965 / 3.067), SIM 0.474 (vs ~0.10), NSS 4.398 (vs 1.484). On the Easy split it reports KLD 1.270 / SIM 0.506 / NSS 4.490. It also generalizes to perspective views on AGD20K (seen split KLD 0.739 / SIM 0.616 / NSS 1.750).

Significance

The panoramic-affordance angle has been underserved β€” most affordance work is per-view, which fails for whole-scene manipulation planning. PanoAffordanceNet is the cleanest 360Β° affordance baseline at the time of writing.

Links

Related pages

← Back to CVPR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally