A carefully curated collection of high-quality tools, libraries, research papers, projects, and tutorials centered around Joint Embedding Predictive Architecture (JEPA) — a self-supervised learning paradigm introduced by Yann LeCun and Meta AI that learns representations by predicting representations of the future from representations of the present, without reconstructing pixels or tokens. This repository serves as a comprehensive, well-organized knowledge hub for researchers and developers exploring the next frontier of self-supervised learning and representation learning.
JEPA represents a fundamental shift in how AI systems learn representations. Unlike traditional generative models that reconstruct inputs, JEPA learns to predict abstract representations of the future state of the world from abstract representations of the present. This approach enables more efficient learning, better generalization, and the ability to handle complex, high-dimensional data without the computational overhead of pixel-level reconstruction.
To keep the community up-to-date with the latest developments, this repository is continuously enriched with newly published JEPA-related papers, real-world use cases, and open-source implementations. From foundational architectures to advanced variants like Hierarchical JEPA (H-JEPA) and applications in vision, language, and multimodal learning, the collection aims to highlight both foundational ideas and emerging best practices.
Note
📢 Announcement: Our paper is now available on SSRN!
Title: A Survey on Joint Embedding Predictive Architectures and World Models
If you find this paper interesting, please consider citing our work. Thank you for your support!
@article{brotee2025survey,
title={A Survey on Joint Embedding Predictive Architectures and World Models},
author={Brotee, Shamyo and Chhetri, Gaurab and Polock, Sazzad Bin Bashar and Bellamkonda, Venkata Surya and Rafe, Amir and Das, Subasish},
journal={Available at SSRN 5772122},
year={2025}
}Whether you are building self-supervised learning systems, researching representation learning, or experimenting with predictive architectures for vision, language, or multimodal tasks, this resource offers a centralized, evolving platform to explore the powerful and expanding universe of JEPA-based systems.
August 25, 2026 at 01:25:56 AM UTC
- PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
- HanoiWorld : A Joint Embedding Predictive Architecture BasedWorld Model for Autonomous Vehicle Controller
- BERT-JEPA: Reorganizing CLS Embeddings for Language-Invariant Semantics
- Value-guided action planning with JEPA world models
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
- Opinion: Learning Intuitive Physics May Require More than Visual Data
- Tokenizing Buildings: A Transformer for Layout Synthesis
- EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants
- Health system learning achieves generalist neuroimaging models
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- POMA-3D: The Point Map Way to 3D Scene Understanding
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- PI-NAIM: Path-Integrated Neural Adaptive Imputation Model
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Multi-Joint Physics-Informed Deep Learning Framework for Time-Efficient Inverse Dynamics
- Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures
- TransactionGPT
- WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
- CSU-PCAST: A Dual-Branch Transformer Framework for medium-range ensemble Precipitation Forecasting
- Improving the Physics of Video Generation with VJEPA-2 Reward Signal
- DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
- Why and How Auxiliary Tasks Improve JEPA Representations
- Valeo Near-Field: a novel dataset for pedestrian intent detection
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
- Joint Embeddings Go Temporal
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
- EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
- MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals
- Predict, Cluster, Refine: A Joint Embedding Predictive Self-Supervised Framework for Graph Representation Learning
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
- Elucidating the Role of Feature Normalization in IJEPA
- TrajFlow: A Generative Framework for Occupancy Density Estimation Using Normalizing Flows
- CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text
- PatchTraj: Unified Time-Frequency Representation Learning via Dynamic Patches for Trajectory Prediction
- Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA Systems
- BadHMP: Backdoor Attack against Human Motion Prediction
- Improving Joint Embedding Predictive Architecture with Diffusion Noise
- From Video to EEG: Adapting Joint Embedding Predictive Architecture to Uncover Visual Concepts in Brain Signal Analysis
- MCST-Mamba: Multivariate Mamba-Based Model for Traffic Prediction
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Conditional Normalizing Flows for Forward and Backward Joint State and Parameter Estimation
- Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
- Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
- VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models
- RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
- Spatiotemporal Semantic V2X Framework for Cooperative Collision Prediction
- WirelessJEPA: A Multi-Antenna Foundation Model using Spatio-temporal Wireless Latent Predictions
- A Latent Space Framework for Modeling Transient Engine Emissions Using Joint Embedding Predictive Architectures
- The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR
- Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
- JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
- Cell-JEPA: Latent Representation Learning for Single-Cell Transcriptomics
- CryoLVM: Self-supervised Learning from Cryo-EM Density Maps with Large Vision Models
- Bayesian Integration of Nonlinear Incomplete Clinical Data
- Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
- MTS-JEPA: Multi-Resolution Joint-Embedding Predictive Architecture for Time-Series Anomaly Prediction
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
- UniSurg: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
- Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency
- Hierarchical JEPA Meets Predictive Remote Control in Beyond 5G Networks
- Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
- Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
- Causal-JEPA: Learning World Models through Object-Level Latent Interventions
- Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces
- Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting
- GOT-JEPA: Generic Object Tracking with Model Adaptation and Occlusion Handling using Joint-Embedding Predictive Architecture
- MeFEm: Medical Face Embedding model
- GenPANIS: A Latent-Variable Generative Framework for Forward and Inverse PDE Problems in Multiphase Media
- JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures
- US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
- PRECTR-V2:Unified Relevance-CTR Framework with Cross-User Preference Mining, Exposure Bias Correction, and LLM-Distilled Encoder Optimization
- Relatron: Automating Relational Machine Learning over Relational Databases
- BiJEPA: Bi-directional Joint Embedding Predictive Architecture for Symmetric Representation Learning
- Improving Diffusion Planners by Self-Supervised Action Gating with Energies
- Escaping The Big Data Paradigm in Self-Supervised Representation Learning
- Hebbian-Oscillatory Co-Learning
- FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
- Representation Learning for Spatiotemporal Physical Systems
- From Video to EEG: Adapting Joint Embedding Predictive Architecture to Uncover Saptiotemporal Dynamics in Brain Signal Analysis
- PASTE: Physics-Aware Scattering Topology Embedding Framework for SAR Object Detection
- Knowledge, Rules and Their Embeddings: Two Paths towards Neuro-Symbolic JEPA
- Laya: A LeJEPA Approach to EEG via Latent Prediction over Reconstruction
- ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning
- LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling
- Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture -- Bridging Predictive and Generative Self-Supervised Learning
- Structured Latent Dynamics in Wireless CSI via Homomorphic World Models
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Interpreting the Synchronization Gap: The Hidden Mechanism Inside Diffusion Transformers
- Probing the Latent World: Emergent Discrete Symbols and Physical Structure in Latent Representations
- SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
- Pretext Matters: An Empirical Study of SSL Methods in Medical Imaging
- A Wireless World Model for AI-Native 6G Networks
- Gaussian Joint Embeddings For Self-Supervised Representation Learning
- JEPA-MSAC: A Joint-Embedding Predictive Architecture for Multimodal Sensing-Assisted Communications
- PI-JEPA: Label-Free Surrogate Pretraining for Coupled Multiphysics Simulation via Operator-Split Latent Prediction
- Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
- Learning General Representation of 12-Lead Electrocardiogram with a Joint-Embedding Predictive Architecture
- From Alignment to Prediction: A Study of Self-Supervised Learning and Predictive Representation Learning
- The Global Neural World Model: Spatially Grounded Discrete Topologies for Action-Conditioned Planning
- A Discordance-Aware Multimodal Framework with Multi-Agent Clinical Reasoning
- JEPAMatch: Geometric Representation Shaping for Semi-Supervised Learning
- Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA
- CasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling
- Why Self-Supervised Encoders Want to Be Normal
- Text-Conditional JEPA for Learning Semantically Rich Visual Representations
- DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring
- AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling
- Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models
- HEPA: A Self-Supervised Horizon-Conditioned Event Predictive Architecture for Time Series
- Multitask Multimodal Fusion with Tabular Foundation Models for Peak and Durability Prediction of Pertussis Booster Response
- Crys-JEPA: Accelerating Crystal Discovery via Embedding Screening and Generative Refinement
- GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
- Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence
- Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction
- Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
- Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
- PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment
- Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives
- SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
- Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation
- ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data
- UWM-JEPA: Predictive World Models That Imagine in Belief Space
- Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion
- Causal-JEPA: Learning World Models through Object-Level Latent Masking
- Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings
- Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models
- HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning
- Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space
- TERRA: Task-Embedded Reasoning and Representation Architecture for Cross-Domain Applications
- UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
- HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
- CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
- DLLM-JEPA: Joint Embedding Predictive Architectures for Masked Diffusion Language Models
- LatentWave: JEPA Pretraining for Wireless Foundation Models
- Predict and Reconstruct: Joint Objectives for Self-Supervised Language Representation Learning
- CF-JEPA: Mask-free forward prediction with asymmetric encoder utilization for time-series representation learning
- iMaC: Translating Actions into Motion and Contact Images for Embodied World Models
- FF-JEPA: Long-Horizon Planning in World Models with Latent Planners
- DALE-CT: Depth-Aware Foundation Models for Computed Tomography
- One Lens, Many Worlds : A Capability-Typed Interface for World-Model Interpretability
- RePAIR: Predictive Self-Supervised Representation Learning in Chess
- Masked and Predictive Self-Supervised Foundation Models for 3D Brain MRI
- Identifiability Without Gaussianity: Symbolic World Models and Near-Infinite Temporal Consistency
- FLaRA: Predicting Future Latent Representations for Accident Anticipation
- Temporal Straightening for Latent Planning
- ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition
- MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation
- Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting
- JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics
- Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow
- SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors
- P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
- Frequency-Aware Self-Supervised Music Representation Learning
- HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
- MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
- A Generalization Theory for JEPA-Based World Models
- Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection
- Fast LeWorldModel
- Domain-Informed Multi-View Self-Distillation for Astronomical Light-Curve Representation Learning with JEPA
- Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture
- A Lightweight Self-Supervised Learning Framework for Multivariate Time Series using Hierarchical-JEPA on ECG Data
- Communication-Aware and Safety-Aware UAV Control via Predictive Latent Models
- AEGIS: A Multi-Task Joint-Embedding Predictive Architecture for Mammography
- Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
- Masked Generative-Contrastive Representation Learning for Cross-Dataset EEG-Based Emotion Recognition
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories
- STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning
- Joint-Embedding Predictive Architecture for Solar PV Panel Fault Classification
- Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method
- Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment
- Learning Subgroup Relations Using Siamese Graph Neural Networks
- JEPA for AI-Native 6G: Predictive Representations and Open Challenges
- MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms
- Contrastive Joint-Embedding Prediction for Representation Learning in Structural MRI
- Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring
- The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models
- Differentiable Cardiac Electrophysiology Simulations for Dynamical State and Parameter Estimation
- A Framework for Early Sepsis Prediction via Self-Supervised (JEPA) and Federated Representation Learning
- Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
- The JEPA Predictor: A Transferable Operator for Occluded Feature Completion
- Scalable and Efficient Joint Spiking Embedding Predictive Architecture for Large-Scale Dynamic Graphs
- Physics-Guided Masked Multi-Task Network for Edge-Friendly Battery Health Diagnostics from Sto-chastically Fragmented Charging Profiles
- Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer
- JEPA-CFM: A Joint Embedding Predictive Architecture-based Channel Foundation Model for Robust Fluid Antenna Systems
- Direct Bethe Free Energy Minimization for Bayesian Neural Networks
- On the Identifiability of Controlled World Models
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Unbiased Open World Regularization for Fair Self-Supervised Learning
- Music-JEPA: Learning a World Model of Sound from Action
- Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations
- τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
- LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
- JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
- One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
- When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment
- Self-supervised DXA representations encode multi-system disease risk, biological aging and heritability
- Branch-JEPA: Finite-Support Predictive Distributions for JEPA World Models
- Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data
- FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
- HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning
- NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
- SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
- Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture
- BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
- SR-JEPA: Learning Predictive Latent State in 3D Scenes
- Discrete energy as an exact label-free training objective for finite-element surrogates
- UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
- The Maunder Model and Catalog: Stellar Rotation, Bimodal Activity, and Magnetic Braking in Kepler Main-Sequence Stars
- JEPA-WAM: Stage-Level Joint-Embedding Prediction for World-Action Models in Robot Manipulation
- CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
- Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency
- StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
- No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models
- Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields
- WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization
- Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
- MIFR: A Modality-Invariant and Fair Representation Framework for Skin Disease Classification
- Orthogonal JEPA: Factorized Predictive States for Latent World Models
- WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
- When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse
- Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework
We welcome contributions to this repository! If you have a resource that you believe should be included, please submit a pull request or open an issue. Contributions can include:
- New libraries or tools related to JEPA.
- Tutorials or guides that help users understand and implement JEPA.
- Research papers that advance the field of JEPA and self-supervised learning.
- Any other resources that you find valuable for the community
- Fork the repository.
- Create a new branch for your changes.
- Make your changes and commit them with a clear message.
- Push your changes to your forked repository.
- Submit a pull request to the main repository.
Before contributing, take a look at the existing resources to avoid duplicates.
This repository is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material, provided you give appropriate credit, link to the license, and indicate if changes were made.