全双工语音模型在被询问时开口,而非在需要时开口
Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
- University of Connecticut(康涅狄格大学)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- X Square Robot
- The University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过构建上下文匹配的独白和10种条件,测试了全双工语音模型在何时开口,发现它们更依赖被呼叫和沉默而非虚假事实或危险,且回应内容缺乏质疑或警告,指出了语音启动和内容理解的差距。
AI中文摘要:
全双工语音模型能够同时听和说,有望实现始终在线的助手。然而,它们还必须决定何时应该说话。人类听众在被呼叫或说话者停顿时会说话,但也会自我选择去纠正错误陈述、补充缺失的词语或警告危险。我们询问全双工模型是否也能做到同样的事情。为了将说话的原因与机会分开,我们构建了上下文匹配的英语独白,其中只有触发话语在主题内变化,根据轮流发言规则定义了10种条件,并压缩词间停顿以限制由沉默创造的机会。在五个模型家族中,被呼叫和沉默是远比虚假事实或危险更可靠的触发因素。在Moshi和PersonaPlex中,帧级文本标记概率在触发结束后的前2秒内,对于虚假事实平均低于中性。停顿或允许打断也不能缩小这一差距。当获得发言权时,Moshi和PersonaPlex回答了大多数直接问题,但非空虚假事实回复中质疑该主张的比例仅为0.14至0.15,而危险回复中警告危险的比例为0.04至0.07。因此,本文指出了在语音启动和响应内容方面的差距。缩小这一差距需要真正的内容理解以及基于此的干预决策。
英文摘要:
Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.