Personal list. With relevant research advancing fast and branching out widely, I'll only add papers meeting my needs hereafter.
-
Grounding Image Matching in 3D with MASt3R [ECCV 2024] [mast3r]
-
Continuous 3D Perception Model with Persistent State [CVPR 2025] [cut3r]
-
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass [CVPR 2025] [fast3r-3d]
-
MUSt3R: Multi-view Network for Stereo 3D Reconstruction [CVPR 2025] [must3r]
-
PE3R: Perception-Efficient 3D Reconstruction [CVPR 2026] [pe3r]
-
VGGT: Visual Geometry Grounded Transformer [CVPR 2025] [vggt]
-
Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors [CVPR 2025] []
-
Matrix3D: Large Photogrammetry Model All-in-One [CVPR 2025] [ml-matrix3d]
-
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion [CVPR 2025] [DiffusionSfM]
-
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory [NeurIPS 2025] [Point3R]
-
π³: Scalable Permutation-Equivariant Visual Geometry Learning [ICLR 2026] [Pi3]
-
StreamVGGT: Streaming 4D Visual Geometry Transformer [ICLR 2026] [StreamVGGT]
-
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer [ICLR 2026] [STream3R]
-
WinT3R: Window-Based Streaming Reconstruction With Camera Token Pool [arXiv 2025] [WinT3R]
-
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization [3DV 2026] [sail-recon]
-
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer [ICLR 2026] [FastVGGT]
-
Faster VGGT with Block-Sparse Global Attention [CVPR 2026] [sparse-vggt]
-
Quantized Visual Geometry Grounded Transformer [ICLR 2026] [QuantVGGT]
-
MapAnything: Universal Feed-Forward Metric 3D Reconstruction [3DV 2026] [map-anything]
-
TTT3R: 3D Reconstruction as Test-Time Training [ICLR 2026] [TTT3R]
-
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting [arXiv 2025] [HunyuanWorld-Mirror]
-
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer [CVPR 2026] [OmniVGGT]
-
Depth Anything 3: Recovering the Visual Space from Any Views [ICLR 2026] [DA3]
-
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction [ICLR 2026] [nova3r]
-
AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend [CVPR 2026] [amb3r]
-
HD-VGGT: High-Resolution Visual Geometry Transformer [arXiv 2026] []
-
DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation [CVPR 2026] [DAGE]
-
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation [CVPR 2026] [CARVE]
-
VGGT-Ω [CVPR 2026] [vggt-omega]
-
GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction [arXiv 2026] [GenRecon]
-
Déjà View: Looping Transformers for Multi-View 3D Reconstruction [arXiv 2026] [dvlt]
-
Surflo: Consistent 3D Surface Flow Model with Global State [arXiv 2026] [surflo]
-
🌐 Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes [ECCV 2026] [argus]
-
Offline Feed-Forward 3D Reconstruction at Scale [arXiv 2026] []
-
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training [CVPR 2026] [ZipMap]
-
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory [arXiv 2026] [LoGeR]
-
S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models [ICME 2026] [S-VGGT]
-
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction [CVPR 2026] [scal3r]
-
LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction [arXiv 2026] [lingbot-map]
-
MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry [CVPR 2026] [MERG3R]
-
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction [arXiv 2026] [HorizonStream]
-
Diversity-aware View Partitioning for Scalable VGGT [ECCV 2026] [DA-VGGT]
-
Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations [arXiv 2026] []
-
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis [CVPR 2026] [nas3r]
-
FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views [CVPRF 2026] [ff3r]
-
PanSt3R: Multi-view Consistent Panoptic Segmentation [ICCV 2025] [panst3r]
-
IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction [ICLR 2026] [IGGT_official]
-
Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images [CVPR 2026] [Uni3R]
-
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation [arXiv 2026] []
-
EPS3D : End-to-End Feed-Forward 3D Panoptic Segmentation [ICML 2026] [EPS3D]
-
SegVGGT: Joint 3D Reconstruction and InstanceSegmentation from Multi-View Images [ECCV 2026] [SegVGGT]
-
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion [ICLR 2025] [monst3r]
-
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos [CVPR 2025] [mega-sam]
-
Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving [arXiv 2024] [Driv3R]
-
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction [ICCV 2025] [Geo4D]
-
ViPE: Video Pose Engine for Geometric 3D Perception [arXiv 2025] [vipe]
-
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction [arXiv 2025] [vggt4d]
-
Efficiently Reconstructing Dynamic Scenes One D4RT at a Time [CVPR 2026] [d4rt]
-
ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation [ICLR 2026] [ReconViaGen]
-
UniRecGen: Unifying Multi-View 3D Reconstruction and Generation [arXiv 2025] [UniRecGen]
-
Stepper: Stepwise Immersive Scene Generation with Multiview Panorama [CVPRF 2026] [stepper]
-
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward [arXiv 2026] [VGGRPO]
-
Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models [arXiv 2026] [geometry-grounded-rfm]
-
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space [arXiv 2026] [PixWorld]
-
GenRec: Knowing Where to Reconstruct and Where to Generate [arXiv 2026] [GenRec]
-
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs [arXiv 2024] [splatt3r]
-
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images [ICLR 2025] [NoPoSplat]
-
PreF3R: Pose-Free Feed-Forward 3D Gaussian Splatting from Variable-length Image Sequence [arXiv 2024] [PreF3R]
-
SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction [CVPR 2025] [SPARS3R]
-
LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias [ICLR 2025] [LVSM]
-
FlowR: Flowing from Sparse to Dense 3D Reconstructions [ICCV 2025] [flowr]
-
RayZer: A Self-supervised Large View Synthesis Model [ICCV 2025] [RayZer]
-
AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views [SIGGRAPH Asia 2025] [AnySplat]
-
VGGT-X: When VGGT Meets Dense Novel View Synthesis [arXiv 2025] [VGGT-X]
-
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting [ICLR 2026] [yonosplat]
-
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training [CVPR 2026] [E-RayZer]
-
Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting [CVPR 2026] [OffTheGrid]
-
Sharp Monocular View Synthesis in Less Than a Second [arXiv 2025] [ml-sharp]
-
UniSHARP: Universal Sharp Monocular View Synthesis [arXiv 2026] [UniSHARP]
-
EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Images [CVPR 2026] [ecosplat-site]
-
From Rays to Projections: Better Inputs for Feed-Forward View Synthesis [CVPR 2026] [pvsm-web]
-
Diff3R: Feed-forward 3D Gaussian Splatting with Uncertainty-aware Differentiable Optimization [arXiv 2026] [diff3r]
-
One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion [arXiv 2026] [One_Shot_Refiner]
-
Repurposing Geometric Foundation Models for Multi-view Diffusion [arXiv 2026] [GLD]
-
UniQueR: Unified Query-based Feedforward 3D Reconstruction [arXiv 2026] []
-
Leveling3D: Leveling Up 3D Reconstruction withFeed-Forward 3D Gaussian Splatting andGeometry-Aware Generation [arXiv 2026] []
-
Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting [ICLR 2026] [lgtm]
-
LagerNVS: Latent Geometry for Fully Neural Real-Time Novel View Synthesis [CVPR 2026] [lagernvs]
-
Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors [CVPR 2026] [PFGS360]
-
FreeScale: Scaling 3D scenes via Certainty-Aware Free-View Generation [CVPR 2026] [FreeScale]
-
C3G: Learning Compact 3D Representations with 2K Gaussians [CVPR 2026] [C3G]
-
ZipSplat: Fewer Gaussians, Better Splats [arXiv 2026] [ZipSplat]
-
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model [arXiv 2026] [AnyRecon]
-
QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction [arXiv 2026] [querysplat]