离散与连续:LALMs中统一音频理解的综合研究
Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs
浏览论文内容
中文总结 AI 辅助
本研究系统比较了大型音频语言模型中连续与离散表示,发现语义约束对音频理解至关重要,且扩展模型规模无法弥补表示信息损失,为未来设计提供了平衡指导。
中文摘要 AI 辅助
大型音频语言模型(LALMs)利用连续特征或离散令牌,然而对于通用音频理解而言,最优的表示范式仍存在争议。现有基准测试往往聚焦于狭窄领域,或在LALM上下文之外评估编码器。为弥补这些空白,我们系统性地评估了语音、声音和音乐中的连续与离散表示。利用我们的UniARC框架,结合从SmolLM2-135M到Llama-3-8B的多种模型规模的双重评估策略,我们分析了数据量、模型容量和计算效率之间的动态关系。我们的结果揭示了语义约束在音频理解令牌化中的关键作用,并表明扩展骨干网络无法弥补音频表示中的信息损失,尤其是在数据受限的任务中。这些发现为未来LALMs在平衡语义密度、保真度和效率方面提供了实用指导。
英文摘要
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To address these gaps, we systematically evaluate continuous and discrete representations across speech, sound and music. Utilizing our UniARC framework with dual evaluation strategies across model scales from SmolLM2-135M to Llama-3-8B, we analyze the dynamic relationships of data volume, model capacity, and computational efficiency. Our results reveal the pivotal role of semantic constraints in tokenization for audio understanding and demonstrate that scaling backbones fail to compensate for information loss in audio representation, especially in data-limited tasks. These findings offer practical guidance for balancing semantic density, fidelity, and efficiency in future LALMs.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。