arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09846cs.GR

DynaConTalk:用于长时程且可控的整体性共语3D动作的小波约束扩散

DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion

  • Tohoku University(东北大学)
  • University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, Yoshifumi Kitamura

AI总结:

针对共语动作生成中的平均化和不可控问题,提出小波约束扩散框架DynaConTalk,在SWT系数空间分离姿态与细节,通过动态门控和条件注入实现长时程可控的整体动作生成。

AI中文摘要:

整体性共语动画在动作表示和语音条件两方面都容易产生平均化。在坐标空间扩散中,缓慢的身体姿态、中频的手势动作以及快速的手部或面部细节被纠缠在一个预测目标中,常常产生低方差、过度平滑的动作。同时,密集的节奏和声学线索在固定多模态融合下可能主导稀疏的内容特定信息。我们提出DynaConTalk,一个用于长时程且可控的整体性共语动作生成的小波约束扩散框架。扩散在平稳小波变换(SWT)系数空间中进行,其时间对齐的频带分离了粗略的姿态演变、手势动作和精细的表情细节。我们的动态门控网络保留HuBERT和说话人身份作为基础,并通过运动状态和噪声感知的残差门控选择性地添加节奏、梅尔频谱和文本特征。注意力池化和学习到的深度路由为每个去噪阶段提供互补条件,而帧分辨率节奏路径保持精确的时间。随后,一个有符号的提案-共识更新将这些条件与演变的运动状态协调。匹配噪声约束注入使用相同的采样接口进行历史延续和局部关键姿势修复,并扩展到参考引导控制。分离的身体-手部和面部去噪器,随后进行逆SWT和姿态驱动的根回归器,产生整体动作。实验评估了生成质量、面部准确性、时间连续性和可控编辑。代码、模型和交互式编辑界面可在该https URL获取。

英文摘要:

Holistic co-speech animation is prone to averaging in both motion representation and speech conditioning. In coordinate-space diffusion, slow body posture, mid-frequency gesture strokes, and fast hand or facial details are entangled in one prediction target, often producing low-variance, over-smoothed motion. Meanwhile, dense rhythmic and acoustic cues can dominate sparse content-specific information under fixed multimodal fusion. We present DynaConTalk, a wavelet-constrained diffusion framework for long-form and controllable holistic co-speech motion generation. Diffusion operates in stationary wavelet transform (SWT) coefficient space, whose temporally aligned bands separate coarse posture evolution, gesture strokes, and fine expressive details. Our dynamic gating network preserves HuBERT and speaker identity as a base and selectively adds rhythm, mel, and transcript features through motion-state- and noise-aware residual gates. Attention pooling and learned depth routing deliver complementary conditions to each denoising stage, while a frame-resolution rhythm path preserves precise timing. A signed proposal-consensus update then reconciles these conditions with the evolving motion state. Matched-noise constraint injection uses the same sampling interface for history continuation and localized keypose repair, and extends to reference-guided control. Separate body-hand and facial denoisers, followed by inverse SWT and a pose-driven root regressor, produce holistic motion. Experiments evaluate generation quality, facial accuracy, temporal continuity, and controllable editing. Code, models, and the interactive editing interface are available at https://github.com/zhuyifeiabcd1/DynaConTalk.

补充信息

↑