Can AI truly edit audio, not just generate it? 🎧 Tencent Hy, in collaboration with SJTU, SII, NTU, TJU, ZODA, PKU, FDU, and other collaborators, introduces MMAE. MMAE--A Massive Multitask Audio Editing Benchmark, is the first comprehensive evaluation benchmark for speech and audio "Banana🍌" Instead of simply requiring the AI to "generate" audio, it demands that the AI understand an existing audio clip and precisely modify it according to natural language instructions—altering what needs to be changed while leaving the rest untouched. Current models show an Exact Match Rate (EMR) below 5%, revealing a major gap in reliable audio editing. MMAE includes: ✅ 2,000 high-fidelity samples from real-world scenarios ✅ 17,741 fine-grained rubric evaluation items ✅ 7 modality settings across sound, music, speech and their mixtures ✅ 6 task complexity from basic modifications to multi-hop reasoning and multi-round editing ✅ 8 operation types across local and global granularities How to use: arXiv: https://t.co/TM81ahH7PZ GitHub: https://t.co/UR1dRUKqMD HuggingFace: https://t.co/1MHR1n3LJn Demo: https://t.co/tz2TVHaCk8
Tencent Hunyuan Introduces Benchmark for Precise AI Audio Editing
Tencent HunyuanTencent Hunyuan, in collaboration with several institutions, introduced MMAE, the first comprehensive benchmark for instruction-based audio editing. This new evaluation reveals that current AI models achieve an Exact Match Rate (EMR) below 5% overall, plummeting to 0% for complex, mixed-modality tasks, highlighting a significant gap in AI's ability to precisely modify existing audio based on natural language instructions.
- Exact Match Rate (EMR) for current models
- Below 5% (0% for complex mixed-modality tasks)
- High-fidelity samples
- 2,000
- Rubric evaluation items
- 17,741
- Audio modalities covered
- 7 (sound, music, speech, and mixtures)
- Task complexity levels
- 6 (basic to multi-hop reasoning and multi-round editing)
- Operation types
- 8 (across local and global granularities)
The MMAE benchmark reveals a significant gap in current AI capabilities. Models achieve an Exact Match Rate (EMR) below 5% on its tasks, falling to 0% for complex, mixed-modality scenarios. This exposes bottlenecks in precise execution, contrasting with advancements in visual domains like MAI-Image-2.5.
MMAE provides a diagnostic roadmap and standardized evaluation for next-generation audio editing systems. The benchmark is openly available via arXiv, GitHub, and HuggingFace, offering resources for researchers and developers. This aligns with the broader shift toward intelligent creation and interactive editing, as seen in developments like ElevenLabs Studio Agent.
Still wondering? A few quick answers below.
Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →





