基于语言和风格参考的灵活运动生成
Flexible Motion Generation from Language and Style References
浏览论文内容
中文总结 AI 辅助
FlexMoGen是一个结合文本提示和风格示例片段进行灵活人体运动合成的框架,通过变分风格编码器和潜在扩散模型实现长序列、多风格生成,无需风格监督,在内容保真与风格反映间达到最佳平衡。
中文摘要 AI 辅助
我们提出了FlexMoGen,一个新颖的框架,用于根据自然语言描述和运动风格参考进行灵活的人体运动合成。文本提示在定义语义内容方面很有效,但通常在捕捉精细的风格细节(如节奏、肢体关节运动和表现力动态)方面存在局限。风格示例片段通过直接传达这些微妙的运动特征来补充文本,使模型能够在保留高层次意图的同时再现所需的风格特征。给定文本提示和风格示例片段,FlexMoGen生成高质量的运动,既保留语义内容又忠实反映目标风格,为用户提供对动画生成过程更大的控制。与先前依赖离散风格标签且无法泛化到长序列或多风格生成的方法不同,FlexMoGen学习了一个无需风格监督的变分风格编码器,并支持长时间、时变、多风格合成。我们的框架在一个统一架构中联合预训练风格编码器和文本到运动的潜在扩散模型,通过轻量级适配模块调节运动风格。它集成了高效的相对位置编码方案,并在风格化和非风格化数据集上训练,从而实现对未见过的文本-风格组合的强大泛化能力。实验表明,FlexMoGen在内容保真度和风格反映之间取得了最佳平衡。
英文摘要
We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style references. Text prompts are effective at defining semantic content, but they are often limited in capturing fine-grained style details such as timing, limb articulation, and expressive dynamics. A style example clip supplements the text by conveying these nuanced motion characteristics directly, enabling the model to preserve high-level intent while reproducing the desired stylistic traits. Given a text prompt and a style example clip, FlexMoGen generates high-quality motions that preserve semantic content while faithfully reflecting the target style, offering users greater control over the animation generation process. Unlike prior methods that rely on discrete style labels and do not generalize to long or multi-style generation, FlexMoGen learns a variational style encoder without style supervision and supports long, time-varying, multi-style synthesis. Our framework jointly pre-trains the style encoder and a text-to-motion latent diffusion model within a unified architecture, modulating motion style through a lightweight adaptation module. It integrates an efficient relative positional encoding scheme and is trained on both stylized and non-stylized datasets, enabling strong generalization to unseen text-style combinations. Experiments show that FlexMoGen achieves the best balance between content fidelity and style reflection.
发表机构
- Brown University(布朗大学)
- Epic Games(Epic Games公司)
- University of California, Davis(加州大学戴维斯分校)
- University of Utah(犹他大学)
机构由 AI 辅助整理,请以论文原文为准。