发表机构
Besimple AI(Besimple AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全双工智能体评估忽视轮内适应的问题,提出Duplex Cue方法,区分听者意图与说话者行为,实验显示真人适应率68.2%远超PersonaPlex的34.8%,强调需同时衡量回应方式与是否继续说话。
AI 中文摘要
全双工评估通常强调智能体是继续说话还是停止说话。这种二元对立无法表达人类日常使用的第三种反应:在继续说话的同时纳入听者刚刚贡献的内容。该贡献可能是一个缺失的词语、一个纠正或一个澄清。我们引入了Duplex Cue,一种对全双工语音智能体中这种“轮内适应”的评估方法。Duplex Cue将听者意图(反馈、协作或打断)与说话者行为区分开来:继续不变、在轮内适应或让出。适应包括确认以及内容修订。在一项使用来自无脚本英语对话的300个人工确认线索的单模型案例研究中,我们比较了录制的真人反应与在重放听者音频时生成的PersonaPlex续接。我们保留了208对,其中进行中的说话者在线索开始时处于活跃状态,且每种条件下都有可评分的反应。在66对协作对中,录制的说话者在68.2%的情况下进行适应,而PersonaPlex为34.8%。该模型在其他情况下继续不变(42.4%)或让出(22.7%)。这些发现表明,评估自然语音交互需要衡量智能体如何回应听者的贡献,以及它是否继续说话。
英文摘要
Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this \emph{in-turn adaptation} in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2\% of cases, compared with 34.8\% for PersonaPlex. The model otherwise continues unchanged (42.4\%) or yields (22.7\%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.
Comments12 pages, 8 tables