arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Puppeteer:基于物体的姿态感知同步言语手势生成

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati

arXiv 2609.00369首次发表:更新:

发表机构

Pickford AI; University of Toronto; Vector Institute(皮克福德人工智能公司; 多伦多大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Puppeteer是一种姿态感知、基于物体的同步言语手势扩散模型,通过因果潜在空间合成物理一致的手势,在SceneGes数据集上的实验显示其生成的手势更多样、时间同步性更好且支持基于物体的合成。

AI 中文摘要

生成在时间上连贯、语义与语音对齐且结合周围物体的同步言语手势仍具挑战性。现有的语音驱动手势模型侧重音频-手势对齐,但未明确考虑姿态约束或周围物体,无法捕捉身体手势与物理空间的内在关联。我们提出Puppeteer,这是一种在因果潜在空间中运行的、姿态感知且基于物体的同步言语手势扩散模型。我们将长手势分解为结构化基元,学习因果变分自动编码器将其编码为时间有序的潜在令牌,每个令牌仅依赖于过去的信息。随后,我们直接在因果潜在空间中执行条件扩散,以语音信号、运动历史、初始姿态参考和物体几何为条件,合成物理一致的手势。这种时间有序的潜在公式支持显式时间控制,并支持手势中间补全、手势完成等任务。为了在现有指标之外更好地评估同步言语手势合成,我们引入了针对该任务定制的新评估指标。我们还创建了SceneGes,这是首个精心整理的用于体现式同步言语手势及对应3D物体的合成3D数据集,可实现基于物体的手势生成。实验表明,Puppeteer生成的手势比现有方法更多样、时间同步性更好,同时支持基于物体的手势合成。

英文摘要

Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑