arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10387eess.AScs.CL

GigaChat音频:时间感知大型音频语言模型

GigaChat Audio: Time-aware Large Audio Language Model

  • SaluteDevices, Russia(SaluteDevices)

机构由 AI 辅助整理,请以论文原文为准。

Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

AI总结:

研究音频条件语言模型长录音时间定位难题,提出时间感知音频语言模型,利用合成监督交织时间标记与音频令牌,在多基准测试中实现高精度,支持相关描述与总结,通过消融实验明确多因素影响,并发布模型权重和数据集。

AI中文摘要:

对于音频条件语言模型而言,在长录音中进行时间定位仍然具有挑战性。我们提出了一种时间感知音频语言模型,它能在长达120分钟的输入上,通过明确的时间戳回答问题。我们的方法利用级联管道的大规模合成监督,将周期性时间标记与连续音频令牌交织在一起。我们的模型在短期和长期基准测试中实现了强大的时间定位准确性,并支持时间锚定的片段描述和总结。大量消融实验研究了时间表示、标记频率、分词和持续时间混合设计如何影响准确性和计算成本。我们发布了模型权重和数据集,以支持对时间感知音频理解的进一步研究。

英文摘要:

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

补充信息

↑