Skip to content

Latest commit

 

History

94 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

HKUST MoE Survey

A Survey on Mixture of Experts in
Large Language Models

Awesome PRs Welcome Last Commit

MoE LLMs Timeline
A chronological overview of several representative Mixture-of-Experts (MoE) models in recent years. The timeline is primarily structured according to the release dates of the models. MoE models located above the arrow are open-source, while those below the arrow are proprietary and closed-source. MoE models from various domains are marked with distinct colors: Natural Language Processing (NLP) in green, Computer Vision in yellow, Multimodal in pink, and Recommender Systems (RecSys) in cyan.

MoE LLMs Timeline
Previous Version: January 2025.

Important

Good news! 🎉 Our survey paper has been successfully accepted by TKDE. 🔥🔥🔥

A curated collection of papers and resources on Mixture of Experts in Large Language Models.

Please refer to our survey "A Survey on Mixture of Experts in Large Language Models" for the detailed contents. Paper page

Please let us know if you discover any mistakes or have suggestions by emailing us: wcai738@connect.hkust-gz.edu.cn

Table of Contents

Taxonomy

MoE LLMs Taxonomy

Paper List (Organized Chronologically and Categorically)

  • Less is MoE: Trimming Experts in Domain-Specialist Language Models, [ArXiv 2026], 2026-6-4

  • LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling, [ArXiv 2026], 2026-6-3

  • UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing, [ArXiv 2026], 2026-6-2

  • PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization, [ICML 2026], 2026-6-2

  • DOT-MoE: Differentiable Optimal Transport for MoEfication, [ICML 2026], 2026-6-1

  • DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts, [ICML 2026], 2026-5-31

  • MESA: Improving MoE Safety Alignment via Decentralized Expertise, [ICML 2026], 2026-5-30

  • How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving, [ArXiv 2026], 2026-5-27

  • VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer, [CVPR 2026], 2026-5-27

  • ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference, [ArXiv 2026], 2026-5-26

  • The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence, [ArXiv 2026], 2026-5-26

  • GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs, [ICML 2026], 2026-5-21

  • DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs, [ArXiv 2026], 2026-5-18

  • ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems, [ArXiv 2026], 2026-5-12

  • MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems, [ISCA 2026], 2026-5-7

  • Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs, [ISCA 2026], 2026-5-7

  • GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking, [AAAI 2026], 2026-5-6

  • SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs, [CVPR 2026], 2026-4-27

  • UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training, [ArXiv 2026], 2026-4-21

  • Design and Behavior of Sparse Mixture-of-Experts Layers in CNN-based Semantic Segmentation, [CVPR 2026], 2026-4-15

  • Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling, [CVPR 2026], 2026-4-15

  • WaveMoE: A Wavelet-Enhanced Mixture-of-Experts Foundation Model for Time Series Forecasting, [ICLR 2026], 2026-4-12

  • Do Domain-specific Experts exist in MoE-based LLMs?, [ArXiv 2026], 2026-4-7

  • The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level, [ICML 2026], 2026-4-2

  • On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models, [CVPR 2026], 2026-3-29

  • MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models, [CVPR 2026], 2026-3-26

  • Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization, [ArXiv 2026], 2026-3-23

  • Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers, [ICML 2026], 2026-3-10

  • Scalable Training of Mixture-of-Experts Models with Megatron Core, [ArXiv 2026], 2026-3-8

  • MoE Lens -- An Expert Is All You Need, [ICLR 2025], 2026-3-6

  • RANGER: Sparsely-Gated Mixture-of-Experts with Adaptive Retrieval Re-ranking for Pathology Report Generation, [CVPR 2026], 2026-3-4

  • LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training, [ASPLOS 2026], 2026-2-12

  • Expert Divergence Learning for MoE-based Language Models, [ICLR 2026], 2026-2-10

  • Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs, [ArXiv 2026], 2026-2-9

  • MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm, [ArXiv 2026], 2026-1-13

  • Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism, [ArXiv 2025], 2025-12-25

  • A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models, [ArXiv 2025], 2025-12-3

  • MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts, [ArXiv 2025], 2025-11-26

  • Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs, [ArXiv 2025], 2025-11-10

  • Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs, [ArXiv 2025], 2025-10-1

  • Towards a Comprehensive Scaling Law of Mixture-of-Experts, [ArXiv 2025], 2025-9-28

  • Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment, [ArXiv 2025], 2025-9-26

  • Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say, [ArXiv 2025], 2025-9-25

  • Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs, [ArXiv 2025], 2025-9-12

  • Steering MoE LLMs via Expert (De)Activation, [ArXiv 2025], 2025-9-11

  • MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention, [ArXiv 2025], 2025-6-16

  • Serving Large Language Models on Huawei CloudMatrix384, [ArXiv 2025], 2025-6-15

  • Ming-Omni: A Unified Multimodal Model for Perception and Generation, [ArXiv 2025], 2025-6-11

  • DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts, [ACL 2025], 2025-6-11

  • MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts, [ACL 2025], 2025-6-9

  • HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts, [ArXiv 2025], 2025-5-30

  • MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production, [ArXiv 2025], 2025-5-16

  • Seed1.5-VL Technical Report, [ArXiv 2025], 2025-5-11

  • MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design, [ICML 2025], 2025-5-9

  • Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs, [ArXiv 2025], 2025-5-7

  • MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core, [ArXiv 2025], 2025-4-21

  • MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism, [ArXiv 2025], 2025-4-3

  • MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation, [AAAI 2025], 2025-3-26

  • Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM, [ArXiv 2025], 2025-3-22

  • Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts, [ICLR 2026], 2025-3-7

  • Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs, [ArXiv 2025], 2025-3-7

  • NetMoE: Accelerating MoE Training through Dynamic Sample Placement, [ICLR 2025], 2025-2-28

  • Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts, [ArXiv 2025], 2025-2-27

  • Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization, [ICLR 2025], 2025-2-26

  • Unraveling the Localized Latents: Learning Stratified Manifold Structures in LLM Embedding Space with Sparse Mixture-of-Experts, [ArXiv 2025], 2025-2-19

  • Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient, [ArXiv 2025], 2025-2-7

  • Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models, [ArXiv 2025], 2025-1-21

  • Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models, [ArXiv 2025], 2025-1-21

  • DeepSeek-V3 Technical Report, [ArXiv 2024], 2024-12-27

  • Qwen2.5 Technical Report, [ArXiv 2024], 2024-12-19

  • A Survey on Inference Optimization Techniques for Mixture of Experts Models, [ArXiv 2024], 2024-12-18

  • LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training, [ArXiv 2024], 2024-11-24

  • MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs, [ASPLOS 2025], 2024-11-18

  • Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent, [ArXiv 2024], 2024-11-4

  • MoE-I2: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition, [EMNLP (Findings) 2024], 2024-11-1

  • MoH: Multi-Head Attention as Mixture-of-Head Attention, [ArXiv 2024], 2024-10-15

  • Upcycling Large Language Models into Mixture of Experts, [ArXiv 2024], 2024-10-10

  • GRIN: GRadient-INformed MoE, [ArXiv 2024], 2024-9-18

  • OLMoE: Open Mixture-of-Experts Language Models, [ArXiv 2024], 2024-9-3

  • Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching, [MICRO 2024], 2024-9-2

  • Mixture of A Million Experts, [ArXiv 2024], 2024-7-4

  • Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs, [ArXiv 2024], 2024-7-1

  • Flextron: Many-in-One Flexible Large Language Model, [ICML 2024], 2024-6-11

  • Demystifying the Compression of Mixture-of-Experts Through a Unified Framework, [ArXiv 2024], 2024-6-4

  • Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models, [ArXiv 2024], 2024-6-3

  • MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models, [DAC 2024], 2024-5-29

  • Yuan 2.0-M32: Mixture of Experts with Attention Router, [ArXiv 2024], 2024-5-28

  • MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability, [ArXiv 2024], 2024-5-23

  • Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models, [ArXiv 2024], 2024-5-23

  • Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast, [ArXiv 2024], 2024-5-23

  • MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models, [ArXiv 2024], 2024-5-19

  • Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts, [ArXiv 2024], 2024-5-18

  • M4oE: A Foundation Model for Medical Multimodal Image Segmentation with Mixture of Experts, [MICCAI 2024], 2024-05-15

  • DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, [ArXiv 2024], 2024-5-7

  • Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training, [ArXiv 2024], 2024-5-6

  • Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping, [ArXiv 2024], 2024-4-30

  • M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation Framework, [SIGIR 2024], 2024-4-29

  • Multi-Head Mixture-of-Experts, [ArXiv 2024], 2024-4-23

  • Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone, [ArXiv 2024], 2024-4-22

  • ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling, [EuroSys 2024], 2024-4-22

  • MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts, [ArXiv 2024], 2024-4-22

  • Intuition-aware Mixture-of-Rank-1-Experts for Parameter Efficient Finetuning, [ArXiv 2024], 2024-4-13

  • JetMoE: Reaching Llama2 Performance with 0.1M Dollars, [ArXiv 2024], 2024-4-11

  • Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models, [ArXiv 2024], 2024-4-8

  • Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts, [ICML 2025], 2024-4-7

  • Mixture-of-Depths: Dynamically allocating compute in transformer-based language models [ArXiv 2024], 2024-4-2

  • MTLoRA: A Low-Rank Adaptation Approach for Efficient Multi-Task Learning [ArXiv 2024], 2024-3-29

  • Jamba: A Hybrid Transformer-Mamba Language Model, [ArXiv 2024], 2024-3-28

  • Scattered Mixture-of-Experts Implementation, [ArXiv 2024], 2024-3-13

  • Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM [ArXiv 2024], 2024-3-12

  • HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts, [ACL 2024], 2024-2-20

  • Higher Layers Need More LoRA Experts [ArXiv 2024], 2024-2-13

  • FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion, [ArXiv 2024], 2024-2-5

  • MoE-LLaVA: Mixture of Experts for Large Vision-Language Models [ArXiv 2024], 2024-1-29

  • OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models [ArXiv 2024], 2024-1-29

  • LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs [ArXiv 2024], 2024-1-29

  • Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference [ArXiv 2024], 2024-1-16

  • MOLE: MIXTURE OF LORA EXPERTS [ICLR 2024], 2024-1-16

  • DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models [ArXiv 2024], 2024-1-11

  • Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning [ArXiv 2023], 2023-12-19

  • LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin [ArXiv 2023], 2023-12-15

  • Mixtral of Experts, [ArXiv 2024], 2023-12-11

  • LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training, [Github 2023], 2023-12

  • Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts [ArXiv 2023], 2023-12-1

  • HOMOE: A Memory-Based and Composition-Aware Framework for Zero-Shot Learning with Hopfield Network and Soft Mixture of Experts, [ArXiv 2023], 2023-11-23

  • Sira: Sparse mixture of low rank adaptation, [ArXiv 2023], 2023-11-15

  • SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models, [MLSys 2024], 2023-10-29

  • When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical Applications, [SIGIR 2024], 2023-10-21

  • Unlocking Emergent Modularity in Large Language Models, [NAACL 2024], 2023-10-17

  • Merging Experts into One: Improving Computational Efficiency of Mixture of Experts, [EMNLP 2023], 2023-10-15

  • Sparse Universal Transformer, [EMNLP 2023], 2023-10-11

  • SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-Prompts, [EMNLP 2023], 2023-10-8

  • FUSING MODELS WITH COMPLEMENTARY EXPERTISE [ICLR 2024], 2023-10-2

  • Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning [ICLR 2024], 2023-9-11

  • EdgeMoE: Fast On-Device Inference of MoE-based Large Language Models [ArXiv 2023], 2023-8-28

  • Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference [ArXiv 2023], 2023-8-23

  • Robust Mixture-of-Expert Training for Convolutional Neural Networks, [ICCV 2023], 2023-8-19

  • Experts Weights Averaging: A New General Training Scheme for Vision Transformers, [ArXiv 2023], 2023-8-11

  • From Sparse to Soft Mixtures of Experts ICLR 2024, 2023-8-2

  • SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization [USENIX ATC 2023], 2023-7-10

  • Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models’ Memories [ACL 2023], 2023-6-8

  • Moduleformer: Learning modular large language models from uncurated data, [ArXiv 2023], 2023-6-7

  • Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural Networks, [ICML 2023], 2023-6-7

  • Soft Merging of Experts with Adaptive Routing, [TMLR 2024], 2023-6-6

  • Brainformers: Trading Simplicity for Efficiency, [ICML 2023], 2023-5-29

  • Emergent Modularity in Pre-trained Transformers, [ACL 2023], 2023-5-28

  • PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts, [ACL 2023], 2023-5-24

  • Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language Models [ICLR 2024], 2023-5-24

  • PipeMoE: Accelerating Mixture-of-Experts through Adaptive Pipelining, [INFOCOM 2023], 2023-5-17

  • MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism [IPDPS 2023], 2023-5-15

  • Optimizing Distributed ML Communication with Fused Computation-Collective Operations, [ArXiv 2023], 2023-5-11

  • FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement [Proc. ACM Manag. Data 2023], 2023-4-8

  • PANGU-Σ: TOWARDS TRILLION PARAMETER LANGUAGE MODEL WITH SPARSE HETEROGENEOUS COMPUTING [ArXiv 2023], 2023-3-20

  • Scaling Vision-Language Models with Sparse Mixture of Experts EMNLP (Findings) 2023, 2023-3-13

  • A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training [ICS 2023], 2023-3-11

  • SPARSE MOE AS THE NEW DROPOUT: SCALING DENSE AND SELF-SLIMMABLE TRANSFORMERS [ICLR 2023], 2023-3-2

  • TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training [NIPS 2022], 2023-2-20

  • PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation, [SOSP 2023], 2023-1-26

  • Mod-Squad: Designing Mixture of Experts As Modular Multi-Task Learners, [CVPR 2023], 2022-12-15

  • Hetu: a highly efficient automatic parallel distributed deep learning system [Sci. China Inf. Sci. 2023], 2022-12

  • MEGABLOCKS: EFFICIENT SPARSE TRAINING WITH MIXTURE-OF-EXPERTS [MLSys 2023], 2022-11-29

  • PAD-Net: An Efficient Framework for Dynamic Networks, [ACL 2023], 2022-11-10

  • Mixture of Attention Heads: Selecting Attention Heads Per Token, [EMNLP 2022], 2022-10-11

  • Sparsity-Constrained Optimal Transport, [ICLR 2023], 2022-9-30

  • A Review of Sparse Expert Models in Deep Learning, [ArXiv 2022], 2022-9-4

  • A Theoretical View on Sparsely Activated Networks [NIPS 2022], 2022-8-8

  • Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models [NIPS 2022], 2022-8-5

  • Towards Understanding Mixture of Experts in Deep Learning, [ArXiv 2022], 2022-8-4

  • No Language Left Behind: Scaling Human-Centered Machine Translation [ArXiv 2022], 2022-7-11

  • Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs [NIPS 2022], 2022-6-9

  • TUTEL: ADAPTIVE MIXTURE-OF-EXPERTS AT SCALE [MLSys 2023], 2022-6-7

  • Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts [NIPS 2022], 2022-6-6

  • Task-Specific Expert Pruning for Sparse Mixture-of-Experts [ArXiv 2022], 2022-6-1

  • Eliciting and Understanding Cross-Task Skills with Task-Level Mixture-of-Experts, [EMNLP 2022], 2022-5-25

  • AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning, [EMNLP 2022], 2022-5-24

  • SE-MoE: A Scalable and Efficient Mixture-of-Experts Distributed Training and Inference System, [ArXiv 2022], 2022-5-20

  • On the Representation Collapse of Sparse Mixture of Experts [NIPS 2022], 2022-4-20

  • Residual Mixture of Experts, [ArXiv 2022], 2022-4-20

  • STABLEMOE: Stable Routing Strategy for Mixture of Experts [ACL 2022], 2022-4-18

  • MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation, [NAACL 2022], 2022-4-15

  • BaGuaLu: Targeting Brain Scale Pretrained Models with over 37 Million Cores [PPoPP 2022], 2022-3-28

  • FasterMoE: Modeling and Optimizing Training of Large-Scale Dynamic Pre-Trained Models [PPoPP 2022], 2022-3-28

  • HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System [ArXiv 2022], 2022-3-28

  • Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models [COLING 2022], 2022-3-2

  • Mixture-of-Experts with Expert Choice Routing [NIPS 2022], 2022-2-18

  • ST-MOE: DESIGNING STABLE AND TRANSFERABLE SPARSE EXPERT MODELS [ArXiv 2022], 2022-2-17

  • UNIFIED SCALING LAWS FOR ROUTED LANGUAGE MODELS [ICML 2022], 2022-2-2

  • One Student Knows All Experts Know: From Sparse to Dense, [ArXiv 2022], 2022-1-26

  • DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale [ICML 2022], 2022-1-14

  • EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate [ArXiv 2021], 2021-12-29

  • Efficient Large Scale Language Modeling with Mixtures of Experts [EMNLP 2022], 2021-12-20

  • GLaM: Efficient Scaling of Language Models with Mixture-of-Experts [ICML 2022], 2021-12-13

  • Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning, [NIPS 2021], 2021.12.6

  • Tricks for Training Sparse Translation Models, [NAACL 2022], 2021-10-15

  • Taming Sparsely Activated Transformer with Stochastic Experts, [ICLR 2022], 2021-10-8

  • MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, [ACL 2022], 2021-10-5

  • Beyond distillation: Task-level mixture-of-experts for efficient inference, [EMNLP 2021], 2021-9-24

  • Scalable and Efficient MoE Training for Multitask Multilingual Models, [ArXiv 2021], 2021-9-22

  • DEMix Layers: Disentangling Domains for Modular Language Modeling, [NAACL 2022], 2021-8-11

  • Go Wider Instead of Deeper [AAAI 2022], 2021-7-25

  • Scaling Vision with Sparse Mixture of Experts [NIPS 2021], 2021-6-10

  • Hash Layers For Large Sparse Models [NIPS 2021], 2021-6-8

  • M6-t: Exploring sparse expert models and beyond, [ArXiv 2021], 2021-5-31

  • BASE Layers: Simplifying Training of Large, Sparse Models [ICML 2021], 2021-5-30

  • FASTMOE: A FAST MIXTURE-OF-EXPERT TRAINING SYSTEM [ArXiv 2021], 2021-5-21

  • CPM-2: Large-scale Cost-effective Pre-trained Language Models, [AI Open 2021], 2021-1-20

  • Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity [ArXiv 2022], 2021-1-11

  • Beyond English-Centric Multilingual Machine Translation, [JMLR 2021], 2020-10-21

  • GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding [ICLR 2021], 2020-6-30

  • Modeling task relationships in multi-task learning with multi-gate mixture-of-experts, [KDD 2018], 2018-7-19

  • OUTRAGEOUSLY LARGE NEURAL NETWORKS: THE SPARSELY-GATED MIXTURE-OF-EXPERTS LAYER [ICLR 2017], 2017-1-23

Contributors

This repository is actively maintained, and we welcome your contributions! If you have any questions about this list of resources, please feel free to contact me at wcai738@connect.hkust-gz.edu.cn.

Star History

Star History Chart

About

[TKDE'25] The official GitHub page for the survey paper "A Survey on Mixture of Experts in Large Language Models".

Resources

Stars

505 stars

Watchers

12 watching

Forks

Releases

Packages

Contributors