发表机构
Xi’an Jiaotong Liverpool University; MiLM Plus, Xiaomi Inc.; University of Oulu; Institute of Acoustics, Chinese Academy of Sciences(西交利物浦大学; 小米公司MiLM Plus; 奥卢大学; 中国科学院声学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoSTALA通过多粒度层次对比学习构建新型训练范式,解决传统ALM难以处理多事件音频序列的问题,为时空音频理解提供强大新框架。
AI 中文摘要
传统音频语言模型(ALM)在实现听觉与文本表征对齐方面已取得显著进展,近期也有针对空间音频的探索。但在日常空间场景中,这类模型仍无法有效处理多事件音频序列。现有方法主要依赖全局听觉与文本特征的粗粒度对比学习,缺乏区分多个顺序事件的分辨率。为克服这些局限,本文提出CoSTALA——一种新型训练范式,从纯全局对齐转向细粒度时空推理。通过构建多粒度层次损失函数体系,实现对时间依赖的显式建模,并成功锚定单个声学事件以保留其语义纯度。大量实验表明,CoSTALA为时空音频理解建立了强大的新框架。
英文摘要
Conventional audio language models (ALMs) have made significant progress in achieving alignment between auditory and textual representations, including recent explorations in spatial audio. However, in daily spatial scenarios, they still cannot effectively process multi-event audio sequences. Current approaches primarily rely on coarse-grained contrastive learning with global auditory and textual features, lacking the resolution to distinguish multiple sequential events. To overcome these limitations, we propose CoSTALA-a novel training paradigm that transitions from purely global alignment to fine-grained spatio-temporal reasoning. By constructing a multi-granularity hierarchical loss function system, we achieve explicit modeling of temporal dependencies, and successfully anchors individual acoustic events to preserve their semantic purity. Extensive experiments demonstrate that CoSTALA significantly establish a powerful new framework for spatio-temporal audio understanding.