arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

控制全双工流式模型中的反馈语

Controlling Backchannels in Streamable Full-duplex Models

Maike Züfle, Peter Polák, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ondřej Klejch

arXiv 2609.29418首次发表:更新:

发表机构

Karlsruhe Institute of Technology; AppTek; Charles University; Sesame AI; University of Edinburgh(卡尔斯鲁厄理工学院; AppTek公司; 查理大学; Sesame AI公司; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种轻量级反馈语头,从全双工模型隐藏状态预测并触发反馈语,跨7B和1B模型规模泛化,生成更频繁且时机更佳的反馈语,质量与真实反馈语相当。

AI 中文摘要

反馈语,例如在对方可能仍在说话时发出的“嗯哼”等简短回应,是自然对话的核心,但全双工口语对话模型很少显式地对其进行建模。我们引入了一个轻量级的反馈语头,它从全双工模型自身的隐藏状态预测反馈语应何时开始。一旦该概率超过可调阈值,就会强制解码一个反馈语。将其附加到7B(PersonaPlex)和1B(F-Actor)模型上,它能够跨规模泛化。探测实验证实隐藏状态能预判真实人类的时机,生成评估显示反馈语更频繁且时机更佳。人类评分者认为生成的反馈语与真实反馈语相当。

英文摘要

Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑