VISTA:视频注入的程式化文本到动画生成
VISTA: Video-Injected Stylized Text-to-Animation
浏览论文内容
中文总结 AI 辅助
VISTA是一个两阶段框架,通过双通道自编码器和掩码自回归扩散,融合文本内容与视频风格生成程式化3D运动,无需配对数据,并在风格识别和内容对齐上表现优异。
中文摘要 AI 辅助
我们提出了VISTA,一个两阶段框架,用于通过融合来自文本提示的结构性内容与来自参考视频的表现性风格来生成程式化的3D人体运动,而无需联合配对的(文本、视频、程式化运动)三元组。一个双通道自编码器首先将运动序列和视频片段映射到一个共享的潜在流形中。然后,一个掩码自回归扩散主干在该流形内运行,通过一个专用的后期融合双自适应层归一化(Dual-AdaLN)路径注入视频衍生的风格,同时保留文本条件的内容结构。一种具有潜在循环一致性的跨批次非配对训练协议,使得能够在语义丰富和风格多样的独立数据集上进行联合学习。作为可控动画合成的概念验证,我们在渲染的动作捕捉参考上验证了VISTA:它在视频条件方法中实现了最高的风格识别准确率,同时保持了有竞争力的内容对齐,其分解的三路无分类器引导在推理时提供了对内容-风格平衡的独立、用户可控的校准。
英文摘要
We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protocol with latent cycle consistency enables joint learning across separate semantically rich and stylistically diverse datasets. As a proof-of-concept for controllable animation synthesis, we validate VISTA on rendered motion-capture references: it achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content-style balance at inference time.
发表机构
- Saarland Informatics Campus(萨尔兰信息学园区)
- German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
- Max Planck Institute for Informatics (MPII)(马克斯·普朗克信息学研究所(MPII))
机构由 AI 辅助整理,请以论文原文为准。