TEMPO:面向大型音频-语言模型的时序锚定多任务后训练
TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models
浏览论文内容
中文总结 AI 辅助
TEMPO是首个处理音频、语音和音乐时间戳任务的统一模型,通过SFT结合GRPO方法,在10K样本评估基准上优于Audio Flamingo Next等最先进模型。
中文摘要 AI 辅助
大型音频-语言模型(LALMs)以片段级描述音频,但无法为其识别的事件、说话人或声音分配时间戳。尽管时间戳对语音识别、密集音频字幕等下游任务至关重要,但仍是大多数LALMs的关键局限。我们提出TEMPO(Temporally-grounded Multi-task Post-training,时序锚定多任务后训练),首个处理音频、语音和音乐时间戳任务的统一模型。我们的核心贡献是基于三项创新构建的监督微调(SFT)阶段:原子时间戳标记、将正弦挂钟编码注入音频帧嵌入的时间感知投影器,以及距离感知高斯损失。我们的训练基于从合成数据到真实数据的课程学习。此外,我们引入了据我们所知首个将强化学习应用于统一音频时间戳的方案,采用带有可验证时序奖励的GRPO,直接优化评估目标。GRPO并非性能提升的主要来源,而是作为SFT检查点之上的精调阶段,提供适度的额外改进。为支持这项工作,我们构建了包含11.9万个样本的训练数据集和包含1万个样本的评估基准,这些数据来自五项任务的既定语料库。在该基准上,TEMPO的性能优于Audio Flamingo Next和Qwen3-Omni这两个明确针对时间戳数据训练的最先进LALMs。实验证实,SFT提供了大部分增益,GRPO则带来稳定但适度的精调效果。
英文摘要
Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for downstream tasks like speech recognition and dense audio captioning, timestamping remains a key limitation of most LALMs. We present TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks. Our core contribution is a supervised fine-tuning (SFT) stage built on three innovations: atomic timestamp tokens, a time-aware projector that injects sinusoidal wall-clock encodings into audio frame embeddings, and a distance-aware Gaussian loss. Our training is based on a synthetic-to-real curriculum. We further introduce, to our knowledge, the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives. Rather than serving as the primary source of performance gains, GRPO acts as a refinement stage on top of the SFT checkpoint, providing modest additional improvements. To support this work, we build a training dataset containing 119K samples and an evaluation benchmark containing 10K samples, drawn from established corpora across five tasks. On this benchmark, TEMPO outperforms Audio Flamingo Next and Qwen3-Omni, two state-of-the-art LALMs explicitly trained on timestamped data. Experiments confirm that SFT delivers most of these gains, with GRPO providing consistent but moderate refinements.
发表机构
- University of Maryland, College Park(马里兰大学帕克分校)
- CNH Industrial India(凯斯纽荷兰工业印度公司)
机构由 AI 辅助整理,请以论文原文为准。