Skip to content

RSS 2026 CLAMP

hwoo.han edited this page Aug 9, 2026 · 2 revisions

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: Manipulation 3 Β· paper #127 Authors: I.-Chun Arthur Liu, Krzysztof Marcin Choromanski, Sandy Huang, Connor Schenck arXiv: 2602.00937 Β· program page

Summary compiled from the arXiv paper (v3, Google DeepMind + USC); all numbers quoted from the paper. Trend context: RSS 2026 survey.

CLAMP pre-training framework and fine-tuning gains (Figure 1 of arXiv 2602.00937, Β© the authors)

Figure 1: three encoders β€” image (multi-view renders from a merged point cloud), action (history a_{tβˆ’1}…a_{tβˆ’H}), and text (object names, positions, task progress) β€” are pulled into a shared embedding space by contrastive learning; the bottom curves show ALOHA Unleashed fine-tuning on Mug-on-Plate with CLAMP (orange, ~0.9 success almost immediately) versus without (blue, plateauing near 0.4 after 1M steps).

Problem

Behavior cloning policies typically reuse pre-trained 2D image representations, which miss the 3D spatial information needed for precise manipulation, and prior robotics pre-training rarely leverages robot actions or 3D perception. It is also unclear which modalities (RGB, point clouds, voxels) best support manipulation pre-training.

Method

CLAMP contrastively pre-trains three encoders on image-text-action triplets: a ViT image encoder over five re-rendered four-channel (depth + XYZ) views from the merged RGB-D point cloud β€” overhead, back-right, front-left, and two dynamic wrist views β€” using STRING relative positional encoding on 3D coordinates so tokens near in 3D space correlate across views (first application of STRING this way); a Transformer action encoder over the action-chunk history; and a CLIP text encoder over privileged simulator text (task description, object names/positions, task-progress integer, used only in pre-training). A sigmoid contrastive loss aligns matched triplets, and a Diffusion Policy (predicting 50Γ—14 noise for the next 50 actions, with ResNet-50 RGB and proprioception inputs) is pre-trained in parallel to initialize fine-tuning. Pre-training uses 553,592 successful episodes across 48 simulated tasks generated by Gemini Robotics 1.5 (~358.5M triplets); fine-tuning uses 4,500 demos per task on 6 unseen tasks.

Results

In simulation, "ALOHA Unleashed with CLAMP" attains 94.0/98.0/51.3/95.3/92.7/95.3% across the six tasks (Can Opener in Caddy, Screwdriver in Caddy, Pen in Container, Mug on Plate, Plate on Rack, Drawer Open), beating ALOHA Unleashed (68.0–80.7%), pretrained variants, ACT, and a Ο€0-architecture VLA baseline; encoders-only pre-training helps, but joint encoder + policy pre-training is markedly stronger (e.g., Mug on Plate 86.0%β†’95.3%). On the real ALOHA robot across five tasks, CLAMP wins or ties everywhere β€” e.g., Open Drawers 7/10 vs 0/10 and Mug on Plate 8/10 vs 3/10 full successes.

Significance

Evidence that action-conditioned 3D contrastive pre-training (plus policy-weight pre-training) is a strong alternative to 2D visual features for bimanual precision manipulation, with wrist-view rendering called out as critical for high-precision tasks. Fits the 3D-representation and pretraining threads in Review-Dexterous-Manipulation and the sim-to-real data-generation angle of Review-Cross-Embodiment.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally