Hansen Jin Lillemark*1
·
Alex Rojas*1
·
Zachary Novack1
·
Runqian Wang2
Yilun Du3
·
Yian Ma1
·
Taylor Berg-Kirkpatrick1
·
Rose Yu1
*Equal contribution 1UC San Diego 2MIT 3Harvard University
Welcome to the codebase for the NeurIPS 2026 paper Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning! EqF proposes to train video denoising generative models without noise level conditions. With an equilibrium field, EqF can leverage adaptive sampling algorithms to improve video generation performance, rather than following hand-designed rigid sampling schedules typically for video generation. The paper's analysis elucidates how noise-unconditional models can perform better than noise-conditional models (e.g. Flow Matching) through avoiding incorrect noise level conditions, how the two classes of models in fact learn the same underlying denoising velocity field, and how the EqF objective incentivizes the model to estimate the noise level of an input internally.
This repository contains the training and inference code for EqF on Minecraft, RealEstate10K (Re10K), and Droid. It also contains the baseline code for standard Diffusion, Flow Matching, and Equilibrium Matching baselines.
Contents:
- Environment Setup
- Quick Start
- Dataset Download
- Model Checkpoint Download
- Inference commands
- Training commands
- Code Walkthrough
- Citation
setup_env.sh creates a conda environment named eqf. The script is meant for Cuda 13:
bash setup_env.sh
# or: bash setup_env.sh --cuda12.4
conda activate eqfTo track runs with wandb, make a copy of
configurations/secrets/my_secrets_template.yaml to
configurations/secrets/my_secrets.yaml and add your API key. Set the wandb entity and project in configurations/config.yaml or override them on the command line.
To get started with inference, download a 256 video subset of the Minecraft validation data, the Minecraft ImageVAE, and the EqF checkpoint:
bash scripts/download_minecraft_mini.shGenerate 32 videos of 300 frames each with:
python -m main shortcode=exp/minecraft/eqf/infer \
+name=minecraft_mini_quickstart dataset.save_dir=data/minecraft/mini \
dataset.num_validation_clips=32Data is expected under data/<dataset>. Download the full datasets with the following commands:
This command downloads the minecraft dataset to data/minecraft/{training,validation}. It may take up to a full day to complete the download:
bash scripts/download_minecraft.sh --workers 11Minecraft training encodes videos online with the released image VAE by default. For faster training, precompute the latents, which can be done in a command described here.
Download and prepare RE10K to data/re10k/{training,validation} with:
bash scripts/download_re10k.sh
python scripts/make_metadata_re10k.pyDownload the raw DROID dataset and caption metadata with the first command. The second command caches the captions with Wan2.2's UMT5 text-prompt embedding model using all visible GPUs:
bash scripts/download_droid.sh
python -m main experiment=cache_prompt_embeds dataset=droid algorithm=wan22_forcing_video +name=cache_droid_prompt_embedsDownload all released checkpoints and dependencies (like pretrained VAEs). Instead of all, you can also pass in {minecraft, re10k, droid} to download just a subset of checkpoints:
bash scripts/download_checkpoints.sh allThe downloaded checkpoint layout is configured in
configurations/ckpt_map/default.yaml, which Hydra selects by default. EqF
checkpoints work for both closed (noise level prediction) and open loop
inference.
Experiments are run using hydra and pytorch lightning, specifying a dataset, algorithm, model family, and hyperparameters. By default, the commands will run on all available GPUs, and are configured for H100 training; adjust device counts and batch sizes for your hardware. Baseline commands are available in the training guide.
# Minecraft EqF model
python -m main shortcode=exp/minecraft/eqf/train_eqf +name=minecraft_eqf
# RE10K EqF model
python -m main shortcode=exp/re10k/eqf/train_pose +name=re10k_eqf
# DROID EqF model
python -m main shortcode=exp/droid/eqf/train +name=droid_eqfTo then train a noise level prediction readout on top of a trained EqF model, there is an additional training command, requiring setting the trained model properly in the checkpoint map. The optimizer state is reset and is trained from the ema weights of the denoising model:
python -m main shortcode=exp/minecraft/eqf/train_noiselevel +name=minecraft_nlp_eqf load=/path/to/minecraft_eqf.ckptEach canonical command evaluates 256 validation videos:
# Minecraft: EqF with NAG, 250 steps, 300 frame videos
python -m main shortcode=exp/minecraft/eqf/infer_nag +name=minecraft_nlp_eqf_nag_infer
# RE10K: EqF with NAG, 50 steps, 37 context + 152 generated frames
python -m main shortcode=exp/re10k/eqf/infer_nag +name=re10k_nlp_eqf_nag_infer
# DROID: budget-adaptive EqF with NAG, 50 steps, 13 context + 36 generated frames
python -m main shortcode=exp/droid/eqf/infer_adaptive +name=droid_eqf_budget_adaptiveThese scripts launch a script to replicate the results in the primary tables with one seed and 256 videos. Use --list to inspect the method indices first, and then --run-indices to select a subset. See the inference guide for more details:
# Fixed ablations table on minecraft (250 steps) and re10k (50 steps)
bash scripts/minecraft_inference.sh
bash scripts/re10k_inference.sh
# Budget-adaptive comparison launchers for EqF NAG and FM at 10-50 steps
bash scripts/adaptive_minecraft_inference.sh
bash scripts/adaptive_re10k_inference.sh
bash scripts/adaptive_droid_inference.shVisualize the dataset and VAE reconstruction using the following commands:
# Minecraft
python main.py experiment=visualize_dataset_and_vae dataset=minecraft \
algorithm/vae=image_vae_minecraft \
+name=viz_minecraft_vae
# Re10k
python main.py experiment=visualize_dataset_and_vae dataset=re10k \
algorithm/vae=wan \
+name=viz_re10k_vae
# Droid
python main.py experiment=visualize_dataset_and_vae dataset=droid \
algorithm/vae=wan22 \
experiment.visualize.bf16=true \
+name=viz_droid_vaeHydra configurations live under configurations/. User-facing recipes are
under configurations/shortcode/; algorithm, dataset, experiment, checkpoint,
and cluster configuration groups can also be selected directly. See
configurations/README.md and the
configuration guide
for conventions.
The full code wiki contains detailed information about the code structure and full training and inference commands.
@misc{lillemark2026equilibriumforcingadaptivevideo,
title={Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning},
author={Hansen Jin Lillemark and Alex Rojas and Zachary Novack and Runqian Wang and Yilun Du and Yian Ma and Taylor Berg-Kirkpatrick and Rose Yu},
year={2026},
eprint={2608.14706},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.14706},
}This repo is forked from FloWM, which is based on Boyuan Chen's research template repo