arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StreamTalk:基于关键姿态锚定的流式伴随语音手势生成

StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi

arXiv 2608.01643首次发表:更新:

发表机构

The University of Tokyo; Alibaba Group; Nanyang Technological University; Renmin University of China(东京大学; 阿里巴巴集团; 南洋理工大学; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

StreamTalk是带周期性生成-检索-精调循环的闭环流式伴随语音手势生成框架,通过关键姿态锚定减少长程漂移,在BEAT2数据集上达到最优FGD指标,实时运行帧率达76 FPS。

AI 中文摘要

实时伴随语音的手势生成需要在语音输入时逐片段生成3D运动片段。现有流式方法为开环结构:每个片段依赖过往上下文,但模型无法检查或修正其运动轨迹,因此小误差会累积并在长序列中产生漂移。我们发现该失败主要由缺乏正向约束而非短片段质量差导致,每个片段末尾的合理关键姿态可提供限制漂移的目标锚点。基于此观察,我们提出StreamTalk,这是一个带周期性生成-检索-精调循环的闭环框架。流式姿态引导生成首先预测粗略片段,从说话人专属运动数据库中检索合理的尾部姿态,再使用该姿态精调片段后进入下一个窗口。训练期间,随机锚点掩码会随机掩码姿态和平移帧,使模型学会从稀疏边界条件中恢复完整运动。感知部件的DiT将手部、身体和平移流分离,以减少全局位移与局部关节运动间的干扰。在BEAT2数据集上,StreamTalk达到了当前最优的FGD指标,相比开环基线降低了长程漂移,且以76 FPS的帧率实时运行。项目页面:this https URL。

英文摘要

Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small errors therefore accumulate and cause drift over long sequences. We observe that this failure is mainly caused by the lack of a forward constraint rather than poor short-clip quality. A plausible key pose at the end of each clip provides a destination anchor that limits drift. Based on this observation, we propose StreamTalk, a closed-loop framework with a periodic generate-retrieve-refine cycle. Streaming Pose-Guided Generation first predicts a coarse clip, retrieves a plausible tail pose from a speaker-specific motion database, and refines the clip using this pose before continuing to the next window. During training, Stochastic Anchor Masking randomly masks pose and translation frames, teaching the model to recover complete motion from sparse boundary conditions. A part-aware DiT separates hand, body, and translation streams to reduce interference between global displacement and local articulation. On BEAT2, StreamTalk achieves state-of-the-art FGD, reduces long-horizon drift relative to open-loop baselines, and runs in real time at 76 FPS. Project page: https://xiangyue-zhang.github.io/StreamTalk/.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑