Reading list on speculative decoding.
Table of Contents
- Bibliography by Venues
- History & Origin
- Draft Models
- Retrieval-based Speculative Decoding
- Draft Tree Construction
- Verification Strategies
- Draft Length Control
- Speculative Decoding + Other Technologies
- Citation
- Other Awesome Lists
-
"SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism" [2025-05] [paper]
-
"Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling" [2025-08] [paper]
-
"FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning" [2025-09] [paper]
-
"Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding" [2025-09] [paper]
-
"Speculative Actions: A Lossless Framework for Faster Agentic Systems" [2025-10] [paper]
-
"Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs" [2025-10] [paper]
-
"Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization" [2025-11] [paper]
-
"Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match" [2025-11] [paper]
-
"Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding" [2026-01] [paper]
-
"Flatter Tokens are More Valuable for Speculative Draft Model Training" [2026-01] [paper]
-
"Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models" [2026-02] [paper]
-
"Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning" [2026-03] [paper]
-
"Speculative Speculative Decoding" [2026-03] [paper]
-
"Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling" [2026-04] [paper]
-
"RepSpec: Structural Re-parameterized Draft Model Training for Speculative Decoding" [2026-04] [paper]
-
"SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications" [2024-11] [paper]
-
"EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization" [2025-02] [paper]
-
"GRIFFIN: Effective Token Alignment for Faster Speculative Decoding" [2025-02] [paper]
-
"Traversal Verification for Speculative Tree Decoding" [2025-05] [paper]
-
"STree: Speculative Tree Decoding for Hybrid State-Space Models" [2025-05] [paper]
-
"DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding" [2025-05] [paper]
-
"MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE" [2025-05] [paper]
-
"List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression" [2025-06] [paper]
-
"Scaling Speculative Decoding with Lookahead Reasoning" [2025-06] [paper]
-
"OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding" [2025-07] [paper]
-
"ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding" [2025-09] [paper]
-
"Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding" [2025-09] [paper]
-
"AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders" [2025-10] [paper]
-
"CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs" [2025-10] [paper]
-
"Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding" [2025-12] [paper]
Main
-
"QSpec: Speculative Decoding with Complementary Quantization Schemes" [2024-10] [paper]
-
"Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation" [2024-11] [paper]
-
"Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference" [2024-12] [paper]
-
"Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification" [2025-05] [paper]
-
"Accelerated Test-Time Scaling with Model-Free Speculative Sampling" [2025-06] [paper]
-
"Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance" [2025-07] [paper]
-
"Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner" [2025-08] [paper]
-
"SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning" [2025-08] [paper]
-
"Speculative Safety-Aware Decoding" [2025-08] [paper]
-
"Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding" [2025-11] [paper]
-
"Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention" [2025-11] [paper]
-
"Cacheback: Speculative Decoding With Nothing But Cache" [2025-11] [paper]
Findings
-
"Speculative Decoding for Multi-Sample Inference" [2025-03] [paper]
-
"MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models" [2025-05] [paper]
-
"Mamba Drafters for Speculative Decoding" [2025-06] [paper]
-
"FractalLLM: Lossless Self-Speculative Decoding with Layer Embedded Self-Compression" [2025-11] [paper]
-
"SpecCoT: Accelerating Chain-of-Thought Reasoning through Speculative Exploration" [2025-11] [paper]
Main
-
"Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling" [2024-08] [paper]
-
"SAM Decoding: Speculative Decoding via Suffix Automaton" [2024-11] [paper]
-
"FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling" [2025-02] [paper]
-
"TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding" [2025-02] [paper]
-
"CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter" [2025-02] [paper]
-
"CLaSp: In-Context Layer Skip for Self-Speculative Decoding" [2025-05] [paper]
-
"A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language Models" [2025-07] [paper]
-
"Faster Speculative Decoding via Effective Draft Decoder with Pruned Candidate Tree" [2025-07] [paper]
-
"SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation" [2025-07] [paper]
Findings
-
"DReSD: Dense Retrieval for Speculative Decoding" [2025-02] [paper]
-
"Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff" [2025-02] [paper]
-
"RASD: Retrieval-Augmented Speculative Decoding" [2025-03] [paper]
-
"Speculative Sampling via Exponential Races" [2025-04] [paper]
-
"Accelerated Diffusion Models via Speculative Sampling" [2025-01] [paper]
-
"Reward-Guided Speculative Decoding for Efficient LLM Reasoning" [2025-01] [paper]
-
"Fast Large Language Model Collaborative Decoding via Speculation" [2025-02] [paper]
-
"Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation" [2025-02] [paper]
-
"Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies" [2025-02] [paper]
-
"QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache" [2025-02] [paper]
-
"RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding" [2025-02] [paper]
-
"Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding" [2025-03] [paper]
-
"SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs" [2025-03] [paper]
-
"Accelerating Large Language Model Reasoning via Speculative Search" [2025-05] [paper]
-
"Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto Speculation" [2025-05] [paper]
-
"BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms" [2025-05] [paper]
-
"Attention-Level Speculation" [2025-07] [paper]
-
"polybasic Speculative Decoding Through a Theoretical Perspective" [2025-07] [paper]
-
"Block Verification Accelerates Speculative Decoding" [2024-03] [paper]
-
"Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference" [2024-05] [paper]
-
"Faster Cascades via Speculative Decoding" [2024-05] [paper]
-
"Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting" [2024-07] [paper]
-
"MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding" [2024-08] [paper]
-
"PEARL: Parallel Speculative Decoding with Adaptive Draft Length" [2024-08] [paper]
-
"Learning Harmonized Representations for Speculative Sampling" [2024-08] [paper]
-
"Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding" [2024-10] [paper]
-
"LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding" [2024-10] [paper]
-
"Mixture of Attentions For Speculative Decoding" [2024-10] [paper]
-
"SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration" [2024-10] [paper]
-
"Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits" [2024-10] [paper]
-
"Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment" [2025-01] [paper]
-
"Towards Optimal Multi-draft Speculative Decoding" [2025-02] [paper]
Main
-
"Decoding Speculative Decoding" [2024-02] [paper]
-
"EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models" [2024-05] [paper]
-
"Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion" [2024-08] [paper]
-
"Constrained Decoding with Speculative Lookaheads" [2024-12] [paper]
Findings
-
"Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding" [2025-02] [paper]
-
"Hierarchical Speculative Decoding with Dynamic Window" [2025-04] [paper]
-
"Cascade Speculative Drafting for Even Faster LLM Inference" [2023-12] [paper]
-
"Sequoia: Scalable and Robust Speculative Decoding" [2024-02] [paper]
-
"Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting" [2024-04] [paper]
-
"Nearest Neighbor Speculative Decoding for LLM Generation and Attribution" [2024-05] [paper]
-
"SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer Devices" [2024-06] [paper]
-
"A Theoretical Perspective for Speculative Decoding Algorithm" [2024-10] [paper]
-
"Inevitable Trade-off between Watermark Strength and Speculative Sampling Efficiency for Language Models" [2024-10] [paper]
-
"Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration" [2024-12] [paper]
Main
-
"Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding" [2024-02] [paper]
-
"Optimized Speculative Sampling for GPU Hardware Accelerators" [2024-06] [paper]
-
"Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters" [2024-06] [paper]
-
"EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees" [2024-06] [paper]
-
"SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding" [2024-11] [paper]
Findings
-
"Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity" [2024-10] [paper]
-
"Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation" [2024-10] [paper]
Main
-
"Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding" [2023-09] [paper]
-
"Speculative Contrastive Decoding" [2023-11] [paper]
-
"LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding" [2024-04] [paper]
Findings
-
"Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding" [2024-01] [paper]
-
"BASS: Batched Attention-optimized Speculative Sampling" [2024-04] [paper]
-
"Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism" [2024-06] [paper]
-
"Graph-Structured Speculative Decoding" [2024-07] [paper]
-
"Online Speculative Decoding" [2023-10] [paper]
-
"Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads" [2024-01] [paper]
-
"EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty" [2024-01] [paper]
-
"Break the Sequential Dependency of LLM Inference Using Lookahead Decoding" [2024-02] [paper]
-
"GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding" [2024-02] [paper]
-
"Accelerated Speculative Sampling Based on Tree Monte Carlo" [2024-07] [paper]
- "DistillSpec: Improving Speculative Decoding via Knowledge Distillation" [2023-10] [paper]
Main
- "REST: Retrieval-Based Speculative Decoding" [2023-11] [paper]
Findings
- "SLiM: Speculative Decoding with Hypothesis Reduction" [2024-06] [paper]
-
"Fast Inference from Transformers via Speculative Decoding" [2022-11] [ICML 2023] [paper]
Experiments on: T5-11B, LaMDA-137B | WMT En-De, CNN/DM
-
"Accelerating Large Language Model Decoding with Speculative Sampling" [2023-02] [paper]
Experiemnts on: Chinchilla-70B | XSum, HumanEval
-
"Accelerating Transformer Inference for Translation via Parallel Decoding" [2023-05] [ACL 2023] [paper]
A block of [PAD] tokens are iteratively refined until no token changes in the block. Applies off-the-shelf to any autoregressive model.
Experiments on: machine translation
-
"PaSS: Parallel Speculative Sampling" [2023-11] [paper]
Learn special tokens (
$L$ lookahead embeddings$[LA]_1, \cdots, [LA]_L$ ) on a small training setExperiments on: LLaMA-7B | Wikipedia, Stack
-
"Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding" [2023-09] [ACL 2024] [paper]
The target LLM selectively skips some of its intermediate layers to generate draft tokens
Experiments on: LLaMA-2-13B/70B, LLaMA-2-13B-Chat, CodeLLaMA-13B | CNN/DM, XSum, HumanEval
-
"Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism" [2024-06] [ACL 2024 Findings] [paper]
Introduces a trainable early-exiting layer on top of the target model's N-th layer hidden states to generate draft tokens
Experiments on: LLaMA-2-13B/70B, LLaMA-2-Chat-70B, Vicuna-13B, CodeLLaMA-13B | GSM8K, XSum, HumanEval, MT-Bench
Draft Model Based on Target Hidden States
-
"Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads" [2024-01] [ICML 2024] [paper]
Train one head (a two-layer FFN with residual connection) for each draft token position
Experiemnts on: Vicuna-7B/13B/33B, Zephyr-7B | MT-Bench
-
"EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty" [2024-01] [ICML 2024] [paper]
The draft model autoregressively processes at the feature (hidden states before LM head) level and then derives tokens using the LM head of the target model
Experiments on: Vicuna-7B/13B/33B, LLaMA2-Chat-7B/13B/70B, Mixtral-8x7B-Instruct | MT-Bench, HumanEval, GSM8K, Alpaca
-
"GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding" [2024-02] [ICML 2024] [paper]
Each layer in the draft model attends to a corresponding layer in the target, counting from the top
Experiments on: Vicuna-7B/13B/33B, Mistral-7B-Instruct-v0.1 | GSM8K, Finance-Alpaca, Spider, CodeSearchNet-Python
-
"Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding" [2024-02] [paper]
A variant of Medusa where each draft head takes output from the previous head as input
Experiments on: Vicuna-7B/13B/33B | MT-Bench
-
"Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding" [2025-03] [paper]
EAGLE for the first two draft tokens, Medusa for the next 5.
Experiments on: Vicuna7B/13B, LLaMA-2-Chat-7B/13B/70B, LLaMA-3-Instruct-8B/70B | MT-Bench, HumanEval, GSM8K, Alpaca, CNN/DM, Natural Questions.
-
"Online Speculative Decoding" [2023-11] [ICML 2024] [paper]
Continuously update the draft model on observed user query data
Experiments on: Vicuna-7B, Flan-T5-3B | Spider, GSM8K, CodeSearchNet-Python, Alpaca-finance
-
"Cascade Speculative Drafting for Even Faster LLM Inference" [2023-12] [NeurIPS 2024] [paper]
- Vertical cascade: in a series of draft models, each model reviews drafts from a smaller one, with the smallest model being a statistical model
- Horizontal cascade: assigns the largest draft model to generate the first draft token, and uses progressively smaller draft models to generate the following tokens (which are less likely to be accepted)
Experiemnts on: Flan-T5, LLaMA-2-Chat-7B | GSM8K, MMLU
-
"Training Domain Draft Models for Speculative Decoding: Best Practices and Insights" [2025-03] [paper]
domain-specific draft models (function calling, biology, Chinese)
Experiments on: LLaMA-3.1-8B
-
"ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts" [2025-03] [paper]
4-bit quantization in MXFP4 datatype as draft model
Experiments on: LLaMA2 7B, Qwen2.5-Coder 7B | HAGRID, MBPP
-
"Automatic Task Detection and Heterogeneous LLM Speculative Decoding" [2025-05] [paper]
Automatically categorizes downstream tasks into different sub-tasks and assigns them to a set of heterogeneous (LoRA-trained) draft models
Experiments on: LLaMA2 13B | Wanjuan 1.0, ChemData700K
-
"Mamba Drafters for Speculative Decoding" [2025-06] [paper]
Experiments on: Pythia 6.9B, Mistral 7B | XSum, CNN/DM, GSM8K, MT-Bench, Alpaca, HumanEval, LongBench
-
"Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding" [2024-02] [EMNLP 2024] [paper]
Extends the last draft token to phrases. Does not require additional training.
Experiments on: Yi-Base-34B, DeepSeek-Coder-Instruct-33B, CodeLLaMA-Instruct-34B, LLaMA-2-Chat-70B | HumanEval, MBPP, GSM8K, CNN/DM, WMT16
-
"SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference" [2024-11] [paper]
N-gram draft model that's built on-the-fly from a suffix tree
Experiments on: LLaMA-3-70B-Instruct | WildChat, Magicoder, Spider, AgenticSQL
-
"SAM Decoding: Speculative Decoding via Suffix Automaton" [2024-11] [paper]
Another n-gram draft model that's built on-the-fly from a suffix tree
Experiments on: Vicuna-7B-v1.3 | MT-Bench, WMT14 De-En, CNN/DM, Natural Question, GSM8K, DPR
-
"Speculative Decoding for Multi-Sample Inference" [2025-03] [paper]
SD method tailored for multi-sample reasoning scenarios, such as self-consistency and Best-of-N sampling.
"for any partial sequence on the path
$i$ , we use its$k$ -token suffix as a query to search for matching prefixes in other paths"Experiments on: Qwen2.5-7B-Instruct, LLaMA-3-8B-Instruct | GSM8K, MATH
-
"GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding" [2024-02] [ICML 2024] [paper]
Expand each draft token to top-$k$ candidates where
$k$ is a piecewise linear function of the top-1 confidence score$p$ , set to 7, 5, 3, 1 for$p$ in (0, 0.3], (0.3, 0.6], (0.6, 0.8], (0.8, 1] respectively.Experiments on: Vicuna-7B/13B/33B | MT-Bench
-
"EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees" [2024-06] [EMNLP 2024] [paper]
In the draft tree, some shallow nodes that are not expanded may have higher values than the deeper expanded nodes. Thus, EAGLE-2 reranks all draft tokens and select the top
$m$ tokens with the highest values.Experiments on: Vicuna-7B/13B, LLaMA2-Chat-7B/13B/70B, LLaMA3-Instruct-8B/70B | MT-Bench, HumanEval, GSM8K, CNN/DM, Natural Questions
-
"HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding" [2025-05] [paper]
HeteroSpec dynamically optimizes computational resource allocation based on linguistic context complexity
- The depth of low-entropy draft paths is increased
- More branches are pruned in low-entropy draft paths to focus on high-likelihood branches.
- Just-in-time graph tracing and compilation are employed to optimize computation graphs.
Experiments on: Vicuna 13B, LLaMA-3.1-Instruct 8B, LLaMA-3.3-Instruct 70B, R1 8B | MT-Bench, HumanEval, GSM8K, Alpaca, CNN/DM.
-
"SpecTr: Fast Speculative Decoding via Optimal Transport" [2023-10] [NeurIPS 2023] [paper]
Introduced OTM - Optimal Transport with Membership cost - and an approximation that's linear in vocabulary size and logarithmic in candidate set size to tackle draft selection when there are multiple drafts
Experiments on: PALM-2-Bison | LM1B
-
"TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding" [2025-02] [paper]
Selects draft tokens for verification based on draft model's output probability
Experiments on: Vicuna-33B-v1.3, LLaMA-3.1-70B/405B-Instruct | ShareGPT, Chatbot Arena, Domain Tough Questions
-
"Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models" [2024-05] [paper]
After generating a draft token, an FFN decides whether or not to continue drafting (input of FFN: top-10 draft model confidence, draft model entropy, token position)
Experiments on: Vicuna-13B-v1.3, Starcoder-15B | CNN/DM, Alpaca, HumanEval, MBPP
-
"Dynamic Depth Decoding: Faster Speculative Decoding for LLMs" [2024-08] [paper]
In the EAGLE framework, use the sum of the probabilities of all the sequences in the beam as a heuristic for whether or not to continue draft generation
Experiments on: Vicuna-7B/13B, LLaMA2-Chat-7B/13B | MT-Bench
-
"Draft Model Knows When to Stop: A Self-Verification Length Policy for Speculative Decoding" [2024-11] [paper]
After generating a draft token, the model decides whether or not to continue draft based on draft model entropy.
Applies off-the-shelf to any autoregressive speculative decoding system without training.
Experiments on: Pythia-6.9B, Vicuna-7B/13B-v1.3, LLaMA-3-70B, Qwen2.5-14B/32B, QwQ | MT-Bench, HumanEval, GSM8K, Alpaca, CNN/DM, Natural Questions, MATH, GPQA, AIME
-
"AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures" [2024-12] [paper]
In the EAGLE framework, train a 3-layer MLP on the penultimate prefix token's input embedding and last hidden states to predict next round's draft length.
Experiments on: Vicuna-7B-v1.3 | MT-Bench, Alpaca, HumanEval, GSM8K, CNN/DM, Natural Questions
-
"SpecServe: Efficient and SLO-Aware Large Language Model Serving with Adaptive Speculative Decoding" [2025-03] [paper]
- Adaptive drafter: incorporates efficiency estimation (based on historical data) into the drafting phase, achieving step-level speculative length control
- Confidence prior verifier: prioritizes the verification of tokens with high acceptance rates, achieving fine-grained request-level speculative length control
- SLO-aware Efficiency estimator: evaluates the efficiency of speculative decoding and achieves SLO (service level objective) awareness
Experiments on: Vicuna-7B-v1.5, Vicuna-33B-v1.3, LLaMA-3.1-70B | MT-Bench, WMT14 De-En, CNN/DM, Natural Questions, GSM8K, DPR
-
"Utility-Driven Speculative Decoding for Mixture-of-Experts" [2025-06] [paper]
Speculative decoding for MoEs. (MoEs break the key assumption in SD that the verification overhead compared to decoding a single token is negligible.)
Experiments on: Mixtral, Phi-3.5, OLMoE, DeepSeek-V1, Qwen1.5 | GSM8K, HumanEval, MT-Bench
-
"Speculative Contrastive Decoding" [2023-11] [ACL 2024 Short] [paper]
Speculative decoding + contrastive decoding to achieve both decoding acceleration and quality improvement
Experiments on: LLaMA-2-70B | WikiText, HumanEval, AlpacaEval, GSM8K
-
"LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding" [2024-04] [ACL 2024] [paper]
- Apply layer dropout (higher dropout rates for later layers) and early exit loss during training
- At inference time, exit at early layers to generate draft tokens, and verify the draft tokens with the ramaining layers.
- Note: this changes target model!
Experiments on: LLaMA-2-7B/13B, LLaMA-3-8B, LLaMA-3.2-1B
-
"DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding" [2025-04] [paper]
DEL, a plug-and-play method that adaptively selects the exit layer and speculation length during inference on top of LayerSkip (2404.16710).
Experiments on: LLaMA-2 7B/13B/70B, LLaMA-3.2 1B, CodeLLaMA 7B/34B | AQuA-RAT, CNN/DM, XSUM, HumanEval
If you refer to this repo, please cite the following paper:
@inproceedings{Zhang2025SVIP,
author = {Ziyin Zhang and
Jiahao Xu and
Tian Liang and
Xingyu Chen and
Zhiwei He and
Rui Wang and
Zhaopeng Tu},
editor = {Christos Christodoulopoulos and
Tanmoy Chakraborty and
Carolyn Rose and
Violet Peng},
title = {Draft Model Knows When to Stop: Self-Verification Speculative Decoding
for Long-Form Generation},
booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural
Language Processing, {EMNLP} 2025, Suzhou, China, November 4-9, 2025},
pages = {16685--16697},
publisher = {Association for Computational Linguistics},
year = {2025},
url = {https://doi.org/10.18653/v1/2025.emnlp-main.844},
doi = {10.18653/V1/2025.EMNLP-MAIN.844}
}
















