Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

14 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination

arXiv Dream Void LLaDA Void

VoidPadding logo

「假作真時真亦假,無為有處有還無。」 -《紅樓夢》太虛幻境

Repeated padding can blur the boundary between emptiness and ending in MDLMs. This ambiguity can cause [EOS] overflow under large-block decoding. VoidPadding decouples the two signals:

  • [VOID] represents padded emptiness during training.
  • [EOS] marks semantic endings.
  • [VOID] generation is banned at inference.

The result is more reliable termination, less wasted generation, and a learned length signal for adaptive canvas expansion.

Method

VoidPadding overview

VoidPadding is evaluated on LLaDA-8B-Instruct and Dream-7B-Instruct. This repository also includes scripts for appendix experiments on LLaDA-8B-Base and Dream-7B-Base.

Results

Fixed-Length Evaluation

Generation length is fixed to 512 and block size is varied over 64, 128, and 512. The Avg. panel reports the mean over GSM8K, MATH500, HumanEval, and MBPP. VoidPadding achieves block-averaged Avg. gains of +7.06/+3.59 on LLaDA and +17.84/+6.95 on Dream over Original/RainbowPadding.

LLaDA fixed-length robustness

Dream fixed-length robustness

Average NFE reduction of VoidPadding relative to Original/RainbowPadding:

Backbone GSM8K MATH500 HumanEval MBPP Avg.
LLaDA 75.0% 13.5% 80.4% 86.4% 63.8%
Dream 69.3% 27.5% 41.3% 84.6% 55.7%

Short-Canvas Expansion

All expansion experiments use an initial response canvas length of 64 and a decoding block size of 32.

For LLaDA, VoidExpansion improves the mean score over Fixed-64 by +11.96, over rho-[EOS] by +5.84, and over DAEDAL by +0.57. It also reduces average NFE from 228.82 to 172.10 relative to DAEDAL, with a 2.12x mean wall-clock speedup.

LLaDA short-canvas expansion accuracy

LLaDA short-canvas expansion efficiency

For Dream, VoidExpansion improves the mean score from 42.23 to 60.37.

Method GSM8K MATH500 HumanEval MBPP Mean
Fixed-64 53.22 22.80 43.90 49.00 42.23
VoidExpansion 79.08 44.00 60.98 57.40 60.37

Quick Start

Create the environment:

bash scripts/setup_env.sh
conda activate void

Download released LoRA adapters:

hf download akpon900/llada-instruct-void \
  --local-dir models/llada-instruct-void

hf download akpon900/dream-instruct-void \
  --local-dir models/dream-instruct-void

Merge them with their base models:

BASE_MODEL="GSAI-ML/LLaDA-8B-Instruct" \
ADAPTER_PATH="models/llada-instruct-void" \
OUTPUT_DIR="outputs/merged/llada-instruct-void" \
bash scripts/merge/merge_llada_lora.sh

BASE_MODEL="Dream-org/Dream-v0-Instruct-7B" \
ADAPTER_PATH="models/dream-instruct-void" \
OUTPUT_DIR="outputs/merged/dream-instruct-void" \
bash scripts/merge/merge_dream_lora.sh

Run evaluation:

bash scripts/eval/fixed_length/llada/run_all_void.sh
bash scripts/eval/expansion/llada/run_all_voidexpansion.sh

bash scripts/eval/fixed_length/dream/run_all_void.sh
bash scripts/eval/expansion/dream/run_all_voidexpansion.sh

Detailed Instructions

Acknowledgements

This repository builds upon and adapts components from:

We thank the authors for releasing their code.

Citation

If you use this repository, please cite the VoidPadding paper.

@misc{voidpadding2026,
  title  = {VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination},
  author = {Chunyu Liu and Zhengyang Fan and Kaisen Yang and Alex Lamb},
  year   = {2026}
}

About

VoidPadding : Let [VOID] handle padding in masked diffusion language models so that [EOS] can focus on semantic termination.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages