arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05126cs.CLcs.MM

语音函数调用:面向大型音频语言模型的语音理解新视角

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

  • Shanghai Jiao Tong University(上海交通大学)
  • Token Foundry, Alibaba Group(阿里巴巴集团Token Foundry)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yuezhang Peng, Yuxin Liu, Changfeng Gao, Zhifu Gao, Qian Chen, Xiangang Li, Xie Chen

AI总结:

该研究提出语音函数调用(SFC)这一新型语义理解视角,构建SFC-Bench数据集并评估LLMs与LALMs性能,经后训练提升LALMs的SFC能力,实验显示SFC优于传统SLU,可大幅提升语义提取准确率。

AI中文摘要:

语音理解(SLU)是面向任务的对话系统的核心组件,也是实现人机无缝交互的关键环节。传统SLU在域内监督微调后可有效提取闭集任务的用户语义,但因规则定义模糊,在开放域任务中利用上下文学习面临重大挑战。本研究提出语音函数调用(SFC)这一新型语义理解视角,通过结构化规则定义优化语义理解,以突破传统闭集SLU的局限。具体而言,我们基于传统SLU数据集整理并扩展了一套语音函数,构建多智能体系统合成SFC-Bench数据集,评估大型语言模型(LLMs)与大型音频语言模型(LALMs)的性能,并通过后训练提升LALMs的SFC能力。实验表明,SFC的表现优于传统SLU,大幅提升了LLMs与LALMs的语义提取准确率。

英文摘要:

Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supervised fine-tuning, it faces significant challenges in leveraging in-context learning for open-domain tasks due to its ambiguous rule definitions. This work proposes Spoken Function Calling (SFC), a novel semantic understanding perspective that optimizes semantic understanding with structured rule definitions, to evolve beyond traditional closed-set SLU. Specifically, we curate and extend a suite of spoken functions based on traditional SLU datasets, construct a multi-agent system to synthesize the SFC-Bench dataset, evaluate the performance of Large Language Models (LLMs) and Large Audio Language Models (LALMs), and enhance the SFC capabilities of LALMs through post-training. Experiments demonstrate that SFC outperforms traditional SLU, substantially enhancing the semantic extraction accuracy for LLMs and LALMs.

补充信息

↑