TEMA:多轮多音频对话中的证据支撑时间问答
TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TEMA通过路由和跨度机制连接事件感知与证据,构建大规模数据集并采用强化学习,提升多轮多音频对话中的时间问答与事件定位能力。
AI中文摘要:
多轮、多音频时间问答要求模型在后续问题、录音切换和历史引用中跟踪目标事件,恢复完整实例及其边界以进行时间计算和比较。我们提出TEMA,它通过Route(指定音频范围)和Span(将相关间隔描述为条件音频字幕)将事件感知与基于证据的问答连接起来。我们构建了包含40,704个对话及每轮证据和答案监督的TEMA-Dialog,以及用于联合评估证据和最终答案的TEMA-Bench。训练结合了时间定位初始化、全对话监督微调以及以完整性为先的仅Span GRPO。在Qwen2.5-Omni和AF-Next上的实验表明,时间问答性能得到提升,特别是在事件定位和跨音频比较方面。仅应用于证据的强化学习进一步提高了间隔恢复和答案准确性。
英文摘要:
Multi-turn, multi-audio temporal question answering requires models to track target events across follow-up questions, recording switches, and historical references, recovering complete instances and their boundaries for temporal calculation and comparison. We propose TEMA, which connects event perception with evidence-based answering through Route, specifying the audio scope, and Span, describing all relevant intervals as conditional audio captions. We construct TEMA-Dialog with 40,704 dialogs and per-turn evidence and answer supervision, and TEMA-Bench for joint evaluation of evidence and final answers. Training combines temporal grounding initialization, full-dialog supervised fine-tuning, and completeness-first Span-only GRPO. Experiments on Qwen2.5-Omni and AF-Next show improved temporal question answering, particularly event localization and cross-audio comparison. Reinforcement learning applied solely to evidence further improves interval recovery and answer accuracy.