聆听、调用与理解:面向大型音频语言模型的技能调用多模态智能体
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
- University of Science and Technology Beijing(北京科技大学)
- Li Auto Inc.(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对工具交互型音频推理问题,开发了SpeechAgent-R智能体,构建了HIU-Corpus与HIU-Bench,经实验验证其在分布内外任务上均较基础模型有显著性能提升。
AI中文摘要:
复杂声学问题可能要求模型执行声学操作、与外部工具交互并对产生的文本或处理后的音频观测结果进行推理,而非直接从固定音频输入中给出答案。我们将此类问题视为工具交互型音频推理问题,并开发了SpeechAgent-R这一音频智能体,它能协调自身的内在多模态理解能力与外部技能和工具。为支持该能力,我们构建了HIU-Corpus,其包含65492条交互轨迹、507.6小时音频,覆盖24项任务、8种技能和9种工具。SpeechAgent-R首先通过基于轨迹的监督微调学习结构化交互行为,随后通过多轮强化学习优化决策。我们还引入了HIU-Bench,用于联合评估任务性能、交互质量及对不同任务设置的泛化能力,该基准包含56项任务的1395个样本,涵盖工具使用和工作流组成存在显著差异的分布内(ID)与分布外(OOD)拆分。SpeechAgent-R在分布内任务上取得84.17的成绩,在分布外任务上取得70.94的成绩,较同一智能体框架下的基础模型分别提升15.40和14.23个百分点。这些结果表明,学习技能与工具的协调能力可提升音频智能体处理多样化任务设置及自适应工具交互的能力。
英文摘要:
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such problems as tool-interactive audio reasoning and develop SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools. To support this capability, we construct HIU-Corpus, comprising 65,492 interaction trajectories and 507.6 hours of audio across 24 tasks, 8 skills and 9 tools. SpeechAgent-R first learns structured interaction behaviors through trajectory-based supervised fine-tuning and then improves its decisions through multi-turn reinforcement learning. We further introduce HIU-Bench to jointly evaluate task performance, interaction quality and generalization to diverse task settings. It contains 1,395 samples across 56 tasks, including in-distribution (ID) and out-of-distribution (OOD) splits with substantial shifts in tool usage and workflow composition. SpeechAgent-R achieves 84.17 on ID tasks and 70.94 on OOD tasks, improving over the base model under the same agent harness by 15.40 and 14.23 points. These results demonstrate that learning skill and tool coordination improves audio agents' ability to handle diverse task settings and adaptive tool interactions.