arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23204cs.ROcs.CLcs.SD

通过真实世界对话机器人中的上下文感知前言生成实现低延迟轮流发言

Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

Yuki Okafuji, Koji Inoue, Yoshiki Ohira

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于LLM的对话系统响应延迟问题,提出两阶段增量框架,通过意图准备检测器和VAP模型生成并确定传递前言回复时机,经现场实验评估三种情况,揭示了不同方式在延迟上的时间权衡。

中文摘要 AI 辅助

基于大语言模型(LLM)的对话系统因仅在最终语音识别后才开始生成回复而存在响应延迟。虽然固定填充语是一种解决方法,但随着时间推移会变得不自然。我们提出了一个两阶段增量框架,将前言回复准备与语音开始解耦。一旦用户意图可预测,意图准备检测器就会触发基于LLM生成简短前言回复。同时,语音活动预测(VAP)模型确定何时传递它。通过在购物中心对路线引导机器人进行现场实验,我们评估了三种情况:无填充语、固定填充语和上下文前言。固定填充语和上下文前言相对于无填充语都显著降低了初始响应延迟。相对于固定填充语,上下文前言的初始响应延迟显著更长,但初始到主要回复的间隔显著更短。探索性评分显示无显著差异。这些结果表明了一种时间权衡。

英文摘要

Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.

发表机构

  • CyberAgent(CyberAgent公司)
  • The University of Osaka(大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑