arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超级明星:面向数字人的流式实时交互智能体

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang

arXiv 2608.24909首次发表:更新:

发表机构

ShanghaiTech University; LIGHTSPEED(上海科技大学; 腾讯光速工作室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对交互式数字人提出流式实时同步言语手势生成任务,构建耦合流式语音响应与在线手势生成的框架,经实验验证其在延迟-质量权衡等方面优于现有基线。

AI 中文摘要

现有的同步言语手势生成方法大多在离线环境中研究,从完整的语音片段合成手势。然而,现实场景中的交互式数字人需要在严格的延迟约束下,仅利用当前可用的响应音频在线生成与语音同步的手势。因此,现有方法不适合实时交互,因为它们要么依赖未来的语音信息,要么会产生大量推理延迟。在本文中,我们针对交互式数字人提出了在线同步言语手势生成任务,并提出了一种实时交互框架,该框架将流式语音响应模块与在线手势生成模块耦合。具体而言,手势生成器被设计为因果多模态自回归模型,可从流式响应语音和运动历史中预测身体运动,无需访问未来语音即可实现低延迟且与语音对齐的手势合成。为支持该设置,我们进一步提出了一种针对虚拟陪伴场景的离线数据合成流程,该流程利用主题和情感感知的主题语料库构建多样化的人与智能体对话,然后基于智能体响应生成同步言语手势。此外,为弥合离线数据构建与在线部署之间的差距,我们建立了一个自进化训练循环,通过将在线交互期间收集的用户反馈纳入数据生成过程,实现对用户偏好的持续适配。大量实验表明,与现有竞争基线相比,我们的框架实现了更优的延迟-质量权衡、更强的语音-运动同步性以及更高的用户偏好。项目页面:this https URL

英文摘要

Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned gesture synthesis without access to future speech. To support this setting, we further propose an offline data synthesis pipeline tailored to virtual companion scenarios, which leverages topic- and emotion-aware subject corpora to construct diverse human-agent dialogues and then generates co-speech gestures conditioned on the agent responses. Moreover, to bridge the gap between offline data construction and online deployment, we establish a self-evolving training loop by incorporating user feedback collected during online interaction into the data generation process, enabling continual adaptation to user preferences. Extensive experiments demonstrate that our framework achieves superior better latency-quality trade-off, stronger speech-motion synchronization, and higher user preference than competitive existing baselines. Project Page: https://super-star-2026.github.io/

CommentsAccepted by ACM Multimedia 2026. Project Page: \url{https://super-star-2026.github.io/}

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑