Tencent Hunyuan Introduces Benchmark for Precise AI Audio Editing

Tencent HunyuanTencent Hunyuan

Tencent Hunyuan, in collaboration with several institutions, introduced MMAE, the first comprehensive benchmark for instruction-based audio editing. This new evaluation reveals that current AI models achieve an Exact Match Rate (EMR) below 5% overall, plummeting to 0% for complex, mixed-modality tasks, highlighting a significant gap in AI's ability to precisely modify existing audio based on natural language instructions.

Tencent Hunyuan, with collaborators, introduced MMAE, a Massive Multitask Audio Editing Benchmark. This benchmark evaluates AI's ability to understand existing audio and precisely modify it based on natural language instructions. It includes 2,000 high-fidelity samples and 17,741 rubric evaluation items (criteria for assessment).
Exact Match Rate (EMR) for current models
Below 5% (0% for complex mixed-modality tasks)
High-fidelity samples
2,000
Rubric evaluation items
17,741
Audio modalities covered
7 (sound, music, speech, and mixtures)
Task complexity levels
6 (basic to multi-hop reasoning and multi-round editing)
Operation types
8 (across local and global granularities)

The MMAE benchmark reveals a significant gap in current AI capabilities. Models achieve an Exact Match Rate (EMR) below 5% on its tasks, falling to 0% for complex, mixed-modality scenarios. This exposes bottlenecks in precise execution, contrasting with advancements in visual domains like MAI-Image-2.5.

MMAE provides a diagnostic roadmap and standardized evaluation for next-generation audio editing systems. The benchmark is openly available via arXiv, GitHub, and HuggingFace, offering resources for researchers and developers. This aligns with the broader shift toward intelligent creation and interactive editing, as seen in developments like ElevenLabs Studio Agent.

Tencent Hy
Tencent Hy
@TencentHunyuan
X

Can AI truly edit audio, not just generate it? 🎧 Tencent Hy, in collaboration with SJTU, SII, NTU, TJU, ZODA, PKU, FDU, and other collaborators, introduces MMAE. MMAE--A Massive Multitask Audio Editing Benchmark, is the first comprehensive evaluation benchmark for speech and audio "Banana🍌" Instead of simply requiring the AI to "generate" audio, it demands that the AI understand an existing audio clip and precisely modify it according to natural language instructions—altering what needs to be changed while leaving the rest untouched. Current models show an Exact Match Rate (EMR) below 5%, revealing a major gap in reliable audio editing. MMAE includes: ✅ 2,000 high-fidelity samples from real-world scenarios ✅ 17,741 fine-grained rubric evaluation items ✅ 7 modality settings across sound, music, speech and their mixtures ✅ 6 task complexity from basic modifications to multi-hop reasoning and multi-round editing ✅ 8 operation types across local and global granularities How to use: arXiv: https://t.co/TM81ahH7PZ GitHub: https://t.co/UR1dRUKqMD HuggingFace: https://t.co/1MHR1n3LJn Demo: https://t.co/tz2TVHaCk8

29retweets185likes
View on X

Still wondering? A few quick answers below.

MMAE is a Massive Multitask Audio Editing Benchmark introduced by Tencent Hunyuan and collaborators. It is designed to evaluate how well AI systems can understand and precisely modify existing audio clips based on natural language instructions.

MMAE is the first comprehensive benchmark for instruction-based audio editing, addressing a gap in AI capabilities. It provides a standardized way to measure progress as AI moves from generating audio to performing precise, targeted edits.

Current AI models show an Exact Match Rate (EMR) below 5% on MMAE tasks. This performance drops to 0% for complex, mixed-modality scenarios, indicating significant challenges in precise execution and structural robustness.

MMAE covers 7 distinct audio modalities, including sound, music, speech, and their mixtures. It includes 6 levels of task complexity, ranging from basic modifications to multi-hop reasoning and multi-round editing, across 8 operation types.

MMAE is openly available for researchers and developers. You can find the associated paper on arXiv, the code on GitHub, and the dataset on HuggingFace.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update