Skip to content

danielqwer/MUGEN

Repository files navigation

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

The official GitHub page of the paper "MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models"

  • Authors: Chih-Kai Yang*, Yun-Shao Tsai*, Yu-Kai Guo†, Ping-Le Tsai†, Yen-Ting Piao‡, Hung-Wei Chen‡, Ting-Lin Hsiao‡, Yun-Man Hsu‡, Ke-Han Lu, Hung-yi Lee (*†‡Equal Contribution)
  • Affiliation: National Taiwan University · NTU AI Center of Research Excellence (NTU AI-CoRE)
  • Paper link: https://arxiv.org/abs/2603.09714

Overview

Overview of MUGEN and the detailed task distribution across the seven evaluation dimensions.

Abstract

TL;DR: We propose MUGEN, a benchmark for multi-audio understanding in LALMs, and show that current models, including proprietary ones, degrade sharply as the number of concurrent audio inputs grows.

While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency (APSC), which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.

Key findings

  • LALMs perform substantially better on semantic tasks than on non-semantic ones, exposing an uneven capability distribution.
  • Accuracy declines steadily as the number of concurrent audio candidates increases — input scaling is a systematic bottleneck.
  • Chain-of-Thought alone fails to resolve auditory perception bottlenecks; Audio-Permutational Self-Consistency (APSC) yields the largest gains, especially when combined with CoT.

News

  • Our paper is now available on arXiv.
  • Our dataset is released on HuggingFace! All 35 tasks are available at MUGEN-Benchmark.

Benchmark

MUGEN contains 9,250 audio clips across 35 tasks organized into 7 evaluation dimensions: Semantics & Pragmatics, Speaker & Demographics, Affective & Paralinguistic State, Temporal Awareness, Acoustic Scene & Event Analysis, Music Analysis, and Compositional Acoustic Reasoning. Each question is formulated as 5-way multiple choice over candidate audio clips; ten tasks additionally include a reference audio for reference-conditioned comparison.

See data/README.md for the data schema and evaluation/README.md for the LLM-judge protocol.

Baselines

  • DeSTA2.5-Audio
  • Qwen2.5-Omni-7B
  • Audio Flamingo 3
  • Voxtral-Mini-3B, Voxtral-Small-24B
  • Phi-4-Multimodal-Instruct
  • Gemini-3-pro (Low / High thinking levels)
  • Cascade baseline: Whisper-large-v3 + Gemini-3-pro (ASR + LLM)

Open-source models are run with vLLM under greedy decoding (Voxtral uses the authors' recommended temperature=0.2, top_p=0.95; Gemini-3-pro uses temperature=1).

Leaderboard

Accuracy (%) on MUGEN across seven dimensions: Semantics & Pragmatics (S&P), Speaker & Demographics (S&D), Affective & Paralinguistic (A&P), Temporal Awareness (TA), Acoustic Scene & Event Analysis (AS&E), Music Analysis (MA), and Compositional Acoustic Reasoning (CA). Overall is the micro-average over all tasks. "Low" / "High" denote the thinking levels of Gemini. Best within each group in bold.

ModelS&PS&DA&PTAAS&EMACAOverall
Open-source LALMs
DeSTA2.5-Audio46.6722.0021.7517.2030.4018.0028.4024.91
Qwen2.5-Omni70.0012.0032.5015.2010.8053.2018.0028.69
Audio Flamingo 325.3322.5015.0021.2023.201.2019.2017.43
Voxtral-Mini-3B59.3325.5030.0020.0022.4024.8023.6027.83
Voxtral-Small-24B64.6727.5032.5022.8022.4025.6016.8028.63
Phi-4-Multimodal-Instruct52.0021.0021.0028.0024.8024.8024.8026.29
Proprietary LALM
Gemini-3-pro (Low)89.3368.0077.0048.0055.6069.6069.2067.66
Gemini-3-pro (High)90.6771.5077.5053.6056.8068.4072.8069.60
Cascaded systems
ASR + LLM (Low)75.3320.0027.0028.8026.4026.0018.0029.09
ASR + LLM (High)82.6716.5027.7528.8026.0026.0022.4030.06

Citation

If you find MUGEN helpful, please consider citing our paper:

@article{yang2026mugen,
    title={MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models},
    author={Yang, Chih-Kai and Tsai, Yun-Shao and Guo, Yu-Kai and Tsai, Ping-Le and Piao, Yen-Ting and Chen, Hung-Wei and Hsiao, Ting-Lin and Hsu, Yun-Man and Lu, Ke-Han and Lee, Hung-yi},
    journal={arXiv preprint arXiv:2603.09714},
    year={2026}
}

About

MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages