发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究全双工语音模型在无脚本对话中相互对话时的话轮转换时序,发现其耦合但滞后于人类,且依赖反应性等待而非话轮结束投射。
AI 中文摘要
全双工语音模型被训练用于与人对话,但它们越来越多地被用于彼此对话,例如在自博弈数据生成、智能体社会以及基于模型的评估中。在这种循环中,没有人来吸收时序错误:每个模型的话轮转换成为另一个模型的输入。我们探究该循环最终会稳定在何种时序上。两个PersonaPlex-7B实例在共享时钟上于无脚本对话中交换音频令牌,并将一种话轮转移规则应用于它们及Switchboard数据集。它们的时序是耦合的:跨对话重新配对说话者会破坏这种耦合。但话轮转换发生得较晚,中位数为400-560毫秒,而人类为137毫秒,且伙伴话轮的最后120毫秒(人类投射在此处放置了十分之一的转移)仅占他们转移的1%。延迟通道的一个方向会使响应一对一地延迟,并使其前的准备时间变为空,这与在感知到结束后的反应性等待一致,而非人类时序所需的话轮结束投射。
英文摘要
Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner's turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.
CommentsThe paper is under review