arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoleBreak:口语对话中长时程角色扮演鲁棒性的基准测试

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu

arXiv 2609.16614首次发表:更新:

发表机构

The University of Hong Kong; Kami AI(香港大学; Kami AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有口语角色扮演基准短时程且以角色为中心的问题,提出RoleBreak基准,包含310个角色和6,688轮对话,评估九种配置,发现系统语义鲁棒性优于声音情感且长时程鲁棒性脆弱,扩大LLM可改善语义但非声音情感。

AI 中文摘要

语音到语音的对话模型越来越支持角色控制,然而现有的口语角色扮演基准大多以角色为中心且为短时程。这留下了问题:口语对话模型能否在长时间交互中维持多样化的角色,尤其是超越预定义的虚构角色。我们引入了RoleBreak,一个用于口语对话中长时程角色扮演鲁棒性的开放基准。RoleBreak包含310个基于角色和以用户为中心的角色,6,688个经过人工验证的对话轮次,以及11,743条细粒度评估标准,其中1,856个轮次带有表达性情感目标以评估声音情感。其场景设计旨在在长时间对话中强调角色一致性、交互质量、安全性和情感。我们评估了涵盖全双工、全模态和级联ASR-LLM-TTS范式的九种配置。我们发现了四个关键模式。首先,当前系统在语义角色遵循方面明显强于声音情感。其次,语义鲁棒性在长时间交互中仍然脆弱:即使是最强的评估系统,平均也仅在10.4和11.6轮后遇到首次角色和安全失败。第三,扩大LLM规模显著提高了语义鲁棒性并延迟了失败,但对声音情感的改善甚微。最后,用户的声音情感会影响角色扮演行为,即使语言内容固定。这些发现突显了口语角色扮演系统中在长时程鲁棒性和声音表现力方面的持续差距。

英文摘要

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters. We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak contains 310 character-based and user-centered roles, 6,688 human-verified dialogue turns, and 11,743 fine-grained evaluation criteria, with 1,856 turns carrying expressive emotion targets for evaluating vocal emotion. Its scenarios are designed to stress role consistency, interaction quality, safety, and affect over extended conversations. We evaluate nine configurations spanning full-duplex, omni-modal, and cascaded ASR--LLM--TTS paradigms. We find four key patterns. First, current systems are substantially stronger at semantic role adherence than at vocal emotion. Second, semantic robustness remains brittle over long interactions: even the strongest evaluated system encounters its first persona and safety failures after only 10.4 and 11.6 turns on average. Third, scaling the LLM substantially improves semantic robustness and delays failure, but yields little improvement in vocal emotion. Finally, user vocal emotion affects role-playing behavior even when linguistic content is fixed. These findings highlight persistent gaps in both long-horizon robustness and vocal expressiveness in spoken role-playing systems.

Comments5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑