arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增强大型音频-语言模型与帧级定位能力以实现细粒度时间感知

Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo, Haoyu Song, Ian McLoughlin

arXiv 2609.15215首次发表:更新:

发表机构

University of Science and Technology of China; Singapore Institute of Technology(中国科学技术大学; 新加坡理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过为大型音频-语言模型增加帧级定位模型,利用语义查询与细粒度音频特征结合,显著提升事件时间定位的精度和可靠性。

AI 中文摘要

大型音频-语言模型(LALMs)已显著推进了通用音频理解,但在细粒度时间感知方面仍存在局限,尤其是在精确事件定位上。现有方法主要通过对LALMs进行后训练,使其预测事件边界作为时间戳标记。然而,这种生成式公式缺乏时间戳预测与细粒度声学证据之间的明确对应关系,限制了时间定位的精度和可靠性。为解决此问题,我们为LALM增加了一个专用的帧级定位模型,同时利用其语义建模能力来表示事件查询。具体而言,冻结的LALM将事件查询与音频作为上下文进行编码,定位模型将这些查询表示与细粒度音频特征相结合,以在帧级别定位目标事件。在多个时间定位基准上的广泛实验表明,与现有方法相比,我们取得了显著且一致的改进。进一步评估显示,定位模型能够提供时间证据以支持下游推理。

英文摘要

Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative formulation lacks explicit correspondence between the timestamp predictions and fine-grained acoustic evidence, limiting the precision and reliability of temporal localization. To address this issue, we augment the LALM with a dedicated frame-level grounding model while leveraging its semantic modeling capability to represent the event query. Specifically, the frozen LALM encodes the event query with audio as context, and the grounding model combines these query representations with fine-grained audio features to localize the target event at the frame level. Extensive experiments across diverse temporal grounding benchmarks demonstrate strong and consistent improvements over existing methods. Further evaluation shows that the grounding model can provide temporal evidence to support downstream reasoning.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑