UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.