arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向更具表达力的口语大语言模型:细粒度意图基准测试与音-词解耦策略优化

Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization

Xiang Lin, Tian-Hao Zhang, Chunfeng Wang, Zhou Pan, Kun Zhan, Liang Li

arXiv 2608.03054首次发表:更新:

AI 中文总结

该研究针对口语情感对话的基准缺失与策略优化问题,构建了ParaIntent基准测试,提出ALPO方法,在情感表达等指标上优于现有模型。

AI 中文摘要

口语情感对话要求模型理解用户的口语输入,并生成在语义上恰当且情感上富有表现力的回应。这一任务颇具挑战性,因为交际意图既可能通过词汇内容明确表达,也可能通过副语言线索更隐晦地传递,而副语言线索可与词汇本身互补或产生分歧。然而,该领域的进展受到两个限制:一是缺乏区分这些意图表达方式的基准测试,二是缺乏同时兼顾回应质量与情感表达的强化学习目标。为解决合适基准测试缺失的问题,我们推出ParaIntent,这是一个包含14个意图类别的中文基准测试,其中明确样本与隐晦样本数量均衡,同时附带涵盖意图完成度、回应质量与情感表达的多维度评估协议。在策略优化方面,现有方法要么为文本和语音使用共享目标,要么仅对单一模态应用强化学习,导致模态特定的学习信号在策略优化中相互纠缠。基于此,我们提出音-词解耦策略优化(Acoustic-Lexical Decoupled Policy Optimization, ALPO),该方法计算独立的文本优势与声学优势,并将其路由到统一 rollout 中的对应文本 token 与语音 token 中。在相同的奖励函数与训练预算下,ALPO在多数自动指标上优于标准GRPO,且在所有微调变体中取得最佳主观结果,在合成测试集与真人录制测试集上的情感表达性能提升尤为显著。

英文摘要

Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is challenging because communicative intent may be stated explicitly in lexical content or conveyed more implicitly through paralinguistic cues, which can complement or diverge from the words themselves. However, two limitations constrain progress in this area: the scarcity of benchmarks that distinguish these intent expressions, and the lack of reinforcement learning objectives that jointly account for response quality and emotional expression. To address the lack of suitable benchmarks, we introduce ParaIntent, a Chinese benchmark comprising 14 intent categories with balanced explicit and implicit samples, together with a multidimensional evaluation protocol covering intent fulfillment, response quality, and emotional expression. For policy optimization, existing approaches either use a shared objective for text and speech or apply reinforcement learning to only one modality, leaving modality-specific learning signals entangled within policy optimization. Motivated by this, we propose Acoustic-Lexical Decoupled Policy Optimization (ALPO), which computes independent textual and acoustic advantages and routes them to the corresponding text and speech tokens within a unified rollout. Under identical reward functions and training budgets, ALPO improves over standard GRPO on most automatic metrics and achieves the best subjective results among the fine-tuned variants, with particularly clear gains in emotional expressiveness on both the synthetic and human-recorded test sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑