GigaChat音频:时间感知大型音频语言模型
GigaChat Audio: Time-aware Large Audio Language Model
- SaluteDevices, Russia(SaluteDevices)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究音频条件语言模型长录音时间定位难题,提出时间感知音频语言模型,利用合成监督交织时间标记与音频令牌,在多基准测试中实现高精度,支持相关描述与总结,通过消融实验明确多因素影响,并发布模型权重和数据集。
AI中文摘要:
对于音频条件语言模型而言,在长录音中进行时间定位仍然具有挑战性。我们提出了一种时间感知音频语言模型,它能在长达120分钟的输入上,通过明确的时间戳回答问题。我们的方法利用级联管道的大规模合成监督,将周期性时间标记与连续音频令牌交织在一起。我们的模型在短期和长期基准测试中实现了强大的时间定位准确性,并支持时间锚定的片段描述和总结。大量消融实验研究了时间表示、标记频率、分词和持续时间混合设计如何影响准确性和计算成本。我们发布了模型权重和数据集,以支持对时间感知音频理解的进一步研究。
英文摘要:
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.