arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03423cs.AI

DuplexSpeechBench-IFEval:评估全双工语音智能体的隐式指令遵循能力

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

Puneet Mathur, Dinesh Manocha

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DuplexSpeechBench-IFEval基准,评估全双工语音智能体的隐式指令遵循能力,发现不同架构的智能体在话语权管理和人设遵循上存在差异,且面临指令推断、执行及冲突解决等挑战。

中文摘要 AI 辅助

全双工语音智能体必须持续决策何时倾听、发出反馈信号、打断对方、处理语音重叠、获取话语权以及让出话语权。现有基准大多通过显式的轮次管理指令测试这些行为,而实际部署的智能体通常由角色或人设配置,需从中推断出合适的对话行为。本文介绍用于评估实时口语交互中隐式指令遵循能力的DuplexSpeechBench-IFEval(DSB-IFEval),该基准包含1038个测试用例,覆盖8种不同的助手角色,评估5种用于指令遵循的条件协议:默认行为、显式行为指令、人设隐含行为、人设-规则组合条件以及指令冲突。我们使用确定性的指令遵循分数(IAS)衡量实时话语权管理,使用大模型评判的人设遵循分数(PAS)衡量人设一致的内容。在6个实时语音系统上,我们发现存在依赖架构的权衡:全双工模型如F-Actor和PersonaPlex对对话行为是明确说明还是需从人设推断更为敏感,在仅人设条件下,其遵循率分别下降9.7%和4.5%;相比之下,GPT-Realtime、MiniCPM-o和Fun-Audio-Chat能强遵循人设一致的内容,但它们的话语权行为在显式和仅人设指令间无适配,且在多项主动行动上仍受限制。我们进一步发现,即使系统能可靠遵循与指定人设冲突的指令,在安全冲突下仍难以覆盖人设要求。这些结果表明,推断角色隐含的行为、在合适的对话时刻执行该行为以及解决相互冲突的指令,仍是全双工语音智能体面临的不同挑战。

英文摘要

Full-duplex voice agents must continuously decide when to speak, listen, backchannel, interrupt, overlap, and yield the conversational floor. Existing benchmarks evaluate these behaviors through explicit turn-management instructions, whereas voice agents are often configured through roles or personas from which appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval), a benchmark for evaluating implicit instruction following in real-time spoken interaction. DSB-IFEval comprises 1,038 test cases derived from 240 controlled conversations spanning eight behaviorally contrastive assistant roles and five conditioning protocols: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona-rule conditioning, and instruction conflict. We measure real-time floor management using the deterministic Instruction Adherence Score (IAS) and persona-consistent response content using the LLM-judged Persona Adherence Score (PAS). Across eleven real-time speech models, we find that executing an explicit floor-management policy does not reliably imply the ability to infer the same policy from a persona. Moreover, even frontier models such as GPT-Live-1 and Gemini-3.8-Live adapt their dialogue language to the assigned persona without consistently translating that persona into the appropriate full-duplex floor-management behavior. Finally, models that successfully resolve benign instruction conflicts often fail when safety-relevant role behavior should override an explicit directive. These results show that inferring role-implied behavior, executing it in real time, and resolving instruction conflicts remain distinct challenges for full-duplex voice agents.

发表机构

  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑