arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00272eess.AScs.SD

当意图迟到:延迟意图揭示下的全双工语音模型基准

When Intent Arrives Late: A Benchmark for Full-Duplex Speech Models under Delayed Intent Revelation

Yang Xiao, Tianyi Peng, Hanyu Meng, Ting Dang

首次发表
浏览论文内容

中文总结 AI 辅助

针对全双工语音模型在延迟意图揭示下的行为,提出LateIntent-Bench配对基准和过早响应率指标,发现多数模型在延迟意图下增加有害参与,插入静音可缓解,凸显评估延迟意图的必要性。

中文摘要 AI 辅助

原生全双工语音模型可以在用户说完之前做出响应,这使得模型的行为既取决于最终话语,也取决于意图定义信息到达的时机。现有基准主要评估轮流说话机制,而安全评估通常假设在生成响应之前已观察到完整请求。我们引入了LateIntent-Bench,一个用于延迟意图揭示的配对基准。一个共享的模糊前缀后面接良性或有害的延续,通过受控停顿来延迟分支变得可区分的时机。我们定义了过早响应率(PRR)来衡量响应开始是否先于意图揭示。对有害和良性参与的联合评估区分了安全选择性的变化与响应性的普遍丧失。在3,136个会话中,四个原生全双工模型在延迟意图下表现出不同的响应时序模式。三个模型在保持良性响应的同时增加了有害参与,而一个模型对两个分支都失去了参与。在意图揭示后插入1.5秒的静音使PRR接近无停顿基线,并大幅减少有害参与的变化。这些结果表明,仅评估完全指定的请求无法捕捉新兴的时序模式,突显了在全双工模型中评估延迟意图揭示的必要性。代码将很快公开发布。

英文摘要

Native full-duplex speech models can respond before a user finishes speaking, making the behavior of the model depend on both the final utterance and when intent-defining information arrives. Existing benchmarks primarily evaluate turn-taking mechanics, whereas safety evaluations typically assume that the complete request is observed prior to response generation. We introduce LateIntent-Bench, a matched-pair benchmark for delayed intent revelation. A shared ambiguous prefix precedes either a benign or a harmful continuation, using a controlled pause to delay when the branches become distinguishable. We define the Premature Response Rate (PRR) to measure whether response onset precedes the revelation of intent. Joint evaluation of harmful and benign engagement distinguishes changes in safety selectivity from general losses in responsiveness. Across 3,136 sessions, four native full-duplex models exhibit distinct response-timing patterns under delayed intent. Three models show increased harmful engagement while maintaining benign responsiveness, whereas one loses engagement with both branches. Inserting a 1.5s silence after intent revelation keeps PRR near the no-pause baseline and substantially reduces changes in harmful engagement. These results demonstrate that evaluating fully specified requests alone fails to capture emerging timing patterns, highlighting the necessity of assessing delayed intent revelation in full-duplex models. The code will be publicly released soon.

发表机构

  • University of Melbourne(墨尔本大学)
  • Duke Kunshan University(昆山杜克大学)
  • UNSW Sydney(悉尼新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

↑