Skip to content

Latest commit

 

History

86 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

End-to-End-3D-Reconstruction-Paper-List

Personal list. With relevant research advancing fast and branching out widely, I'll only add papers meeting my needs hereafter.



3D Reconstruction

  • DUSt3R: Geometric 3D Vision Made Easy [CVPR 2024] [dust3r]

  • Grounding Image Matching in 3D with MASt3R [ECCV 2024] [mast3r]

  • 3D Reconstruction with Spatial Memory [3DV 2025] [spann3r]

  • Continuous 3D Perception Model with Persistent State [CVPR 2025] [cut3r]

  • Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass [CVPR 2025] [fast3r-3d]

  • MUSt3R: Multi-view Network for Stereo 3D Reconstruction [CVPR 2025] [must3r]

  • PE3R: Perception-Efficient 3D Reconstruction [CVPR 2026] [pe3r]

  • VGGT: Visual Geometry Grounded Transformer [CVPR 2025] [vggt]

  • Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors [CVPR 2025] []

  • Matrix3D: Large Photogrammetry Model All-in-One [CVPR 2025] [ml-matrix3d]

  • DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion [CVPR 2025] [DiffusionSfM]

  • Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory [NeurIPS 2025] [Point3R]

  • π³: Scalable Permutation-Equivariant Visual Geometry Learning [ICLR 2026] [Pi3]

  • StreamVGGT: Streaming 4D Visual Geometry Transformer [ICLR 2026] [StreamVGGT]

  • STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer [ICLR 2026] [STream3R]

  • WinT3R: Window-Based Streaming Reconstruction With Camera Token Pool [arXiv 2025] [WinT3R]

  • SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization [3DV 2026] [sail-recon]

  • FastVGGT: Training-Free Acceleration of Visual Geometry Transformer [ICLR 2026] [FastVGGT]

  • Faster VGGT with Block-Sparse Global Attention [CVPR 2026] [sparse-vggt]

  • Quantized Visual Geometry Grounded Transformer [ICLR 2026] [QuantVGGT]

  • MapAnything: Universal Feed-Forward Metric 3D Reconstruction [3DV 2026] [map-anything]

  • TTT3R: 3D Reconstruction as Test-Time Training [ICLR 2026] [TTT3R]

  • WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting [arXiv 2025] [HunyuanWorld-Mirror]

  • OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer [CVPR 2026] [OmniVGGT]

  • Depth Anything 3: Recovering the Visual Space from Any Views [ICLR 2026] [DA3]

  • NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction [ICLR 2026] [nova3r]

  • AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend [CVPR 2026] [amb3r]

  • HD-VGGT: High-Resolution Visual Geometry Transformer [arXiv 2026] []

  • DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation [CVPR 2026] [DAGE]

  • Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation [CVPR 2026] [CARVE]

  • VGGT-Ω [CVPR 2026] [vggt-omega]

  • GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction [arXiv 2026] [GenRecon]

  • Déjà View: Looping Transformers for Multi-View 3D Reconstruction [arXiv 2026] [dvlt]

  • Surflo: Consistent 3D Surface Flow Model with Global State [arXiv 2026] [surflo]

  • 🌐 Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes [ECCV 2026] [argus]

Scalable

  • Offline Feed-Forward 3D Reconstruction at Scale [arXiv 2026] []

  • ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training [CVPR 2026] [ZipMap]

  • LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory [arXiv 2026] [LoGeR]

  • S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models [ICME 2026] [S-VGGT]

  • Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction [CVPR 2026] [scal3r]

  • LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction [arXiv 2026] [lingbot-map]

  • MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry [CVPR 2026] [MERG3R]

  • HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction [arXiv 2026] [HorizonStream]

  • Diversity-aware View Partitioning for Scalable VGGT [ECCV 2026] [DA-VGGT]

Self-Supervised

  • Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations [arXiv 2026] []

  • From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis [CVPR 2026] [nas3r]

  • FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views [CVPRF 2026] [ff3r]

Semantic

  • PanSt3R: Multi-view Consistent Panoptic Segmentation [ICCV 2025] [panst3r]

  • IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction [ICLR 2026] [IGGT_official]

  • Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images [CVPR 2026] [Uni3R]

  • VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation [arXiv 2026] []

  • EPS3D : End-to-End Feed-Forward 3D Panoptic Segmentation [ICML 2026] [EPS3D]

  • SegVGGT: Joint 3D Reconstruction and InstanceSegmentation from Multi-View Images [ECCV 2026] [SegVGGT]

Dynamic

  • MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion [ICLR 2025] [monst3r]

  • MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos [CVPR 2025] [mega-sam]

  • Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving [arXiv 2024] [Driv3R]

  • Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction [ICCV 2025] [Geo4D]

  • ViPE: Video Pose Engine for Geometric 3D Perception [arXiv 2025] [vipe]

  • VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction [arXiv 2025] [vggt4d]

  • Efficiently Reconstructing Dynamic Scenes One D4RT at a Time [CVPR 2026] [d4rt]

Generation

Novel View Synthesis

  • Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs [arXiv 2024] [splatt3r]

  • No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images [ICLR 2025] [NoPoSplat]

  • PreF3R: Pose-Free Feed-Forward 3D Gaussian Splatting from Variable-length Image Sequence [arXiv 2024] [PreF3R]

  • SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction [CVPR 2025] [SPARS3R]

  • LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias [ICLR 2025] [LVSM]

  • FlowR: Flowing from Sparse to Dense 3D Reconstructions [ICCV 2025] [flowr]

  • RayZer: A Self-supervised Large View Synthesis Model [ICCV 2025] [RayZer]

  • AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views [SIGGRAPH Asia 2025] [AnySplat]

  • VGGT-X: When VGGT Meets Dense Novel View Synthesis [arXiv 2025] [VGGT-X]

  • YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting [ICLR 2026] [yonosplat]

  • E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training [CVPR 2026] [E-RayZer]

  • Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting [CVPR 2026] [OffTheGrid]

  • Sharp Monocular View Synthesis in Less Than a Second [arXiv 2025] [ml-sharp]

  • UniSHARP: Universal Sharp Monocular View Synthesis [arXiv 2026] [UniSHARP]

  • EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Images [CVPR 2026] [ecosplat-site]

  • From Rays to Projections: Better Inputs for Feed-Forward View Synthesis [CVPR 2026] [pvsm-web]

  • Diff3R: Feed-forward 3D Gaussian Splatting with Uncertainty-aware Differentiable Optimization [arXiv 2026] [diff3r]

  • One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion [arXiv 2026] [One_Shot_Refiner]

  • Repurposing Geometric Foundation Models for Multi-view Diffusion [arXiv 2026] [GLD]

  • UniQueR: Unified Query-based Feedforward 3D Reconstruction [arXiv 2026] []

  • Leveling3D: Leveling Up 3D Reconstruction withFeed-Forward 3D Gaussian Splatting andGeometry-Aware Generation [arXiv 2026] []

  • Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting [ICLR 2026] [lgtm]

  • LagerNVS: Latent Geometry for Fully Neural Real-Time Novel View Synthesis [CVPR 2026] [lagernvs]

  • Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors [CVPR 2026] [PFGS360]

  • FreeScale: Scaling 3D scenes via Certainty-Aware Free-View Generation [CVPR 2026] [FreeScale]

  • C3G: Learning Compact 3D Representations with 2K Gaussians [CVPR 2026] [C3G]

  • ZipSplat: Fewer Gaussians, Better Splats [arXiv 2026] [ZipSplat]

  • AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model [arXiv 2026] [AnyRecon]

  • QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction [arXiv 2026] [querysplat]

About

multi-view, end2end, 3D Reconstruction

Resources

Stars

51 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors