arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07631cs.SDcs.AIcs.MM

PACE:面向基于大语言模型的全双工语音对话的播放对齐上下文引擎

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

Shibo Wang, Zicheng Zhang, Libo Wang, Junfeng Ma

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对全双工语音对话的生成上下文错位问题,提出了中间件层PACE,通过将模型上下文锚定到客户端播放边界解决该问题,在GCM-Bench等数据集上验证了其有效性。

中文摘要 AI 辅助

基于大语言模型(LLM)的全双工语音服务允许用户在助手响应时说话。由于服务器生成输出和推进对话状态的速度快于客户端播放的速度,后续的用户语音可能会基于用户从未听到的内容被解读,我们将这种故障称为生成上下文错位(Generative Context Mis-anchoring,GCM)。为解决GCM问题,我们提出了PACE,这是一种与提供商无关的中间件层,它将面向模型的上下文锚定到客户端播放边界,该边界是用户本可以听到的内容的系统可观测代理。在中断后,PACE会修复此上下文,以排除从未播放的助手内容,同时在异构语音运行时中保持低延迟生成。我们使用黑盒语音模型,在基于浏览器的实时语音助手中端到端实现了PACE的纯音频投影路径,且未修改模型服务。我们还构建了GCM-Bench,这是一个包含108个播放相关指称锚定案例的新型受控基准数据集。在GCM-Bench上,与仅取消的基线相比,PACE将指称锚定准确率从25.0%提高到96.3%。在200个全双工基准v1(Full-Duplex-Bench v1)中断样本上,它保持了中断响应质量。这些结果表明,将面向模型的上下文锚定到实际播放内容是维持全双工语音对话一致性的实用方法。

英文摘要

LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call this failure Generative Context Mis-anchoring (GCM). To address GCM issues, we present PACE, a provider-independent middleware layer that anchors model-facing context to the client playback boundary, a system-observable proxy for what the user could have heard. After an interruption, PACE repairs this context to exclude assistant content that never reached playback, while preserving low-latency generation across heterogeneous voice runtimes. We implement PACE's audio-only projection path end to end in a browser-based realtime voice assistant using a black-box speech model, without modifying the model service. We also construct GCM-Bench, a new controlled benchmark dataset of 108 playback-relative referent-anchoring cases. On GCM-Bench, PACE raises Referent Anchoring Accuracy from 25.0% to 96.3% over a cancellation-only baseline. On 200 Full-Duplex-Bench v1 interruption samples, it preserves interruption response quality. These results show that grounding model-facing context in actual playback is a practical way to maintain consistency in full-duplex voice dialogue.

发表机构

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑