发表机构
University of Rochester; University of Tokyo; University of California Santa Cruz; University of California Los Angeles; Meta(罗切斯特大学; 东京大学; 加州大学圣克鲁兹分校; 加州大学洛杉矶分校; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GestureFAR提出流自回归框架,通过连续潜变量与流匹配头生成流式共语手势,结合头部流蒸馏降低延迟,在BEAT2上实现质量与延迟的显著优化。
AI 中文摘要
从流式语音中生成自然的共语手势对于具身对话代理至关重要,此时必须在用户仍在说话的同时生成动作。最近的流式手势系统通过自回归离散运动令牌实现了在线生成,但这种设计将高维连续运动压缩到有限的码本中,可能限制生成手势的真实感和多样性。为了同时保持因果性和连续表现力,我们提出了GestureFAR,一种用于流式共语手势生成的流自回归框架。首先,GestureFAR对因果连续运动潜变量进行自回归,使用Transformer建模流式音频-运动上下文,并通过逐令牌的流匹配头从连续分布中采样下一个潜变量。其次,我们引入了一种仅头部的流蒸馏策略,冻结因果主干,并使用一致性和分布匹配目标将多步逐令牌流头蒸馏为单次网络评估。这保持了模型的令牌因果性,同时消除了实时交互的主要延迟瓶颈。在BEAT2上的实验表明,GestureFAR显著改善了流式方法中的质量-延迟权衡,在保持强大手势质量的同时实现了实时令牌因果生成。项目页面:此https URL
英文摘要
Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR