通过超网络驱动的低秩适应生成风格化文本到动作生成
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
- Visual Media Lab, KAIST(韩国庆熙大学视觉媒体实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出轻量级风格条件框架,通过超网络生成LoRA参数动态调节预训练扩散模型,实现高效且通用的风格化动作生成。
AI中文摘要:
文本驱动的运动扩散模型能够生成逼真的人体动作,但文本单独往往难以表达动作的细粒度风格。最近的方法通过在预训练文本驱动扩散模型上附加风格注入机制来解决这一挑战。然而,现有风格化方法要么需要对现有模型进行特定风格的微调,要么依赖于重型ControlNet架构,限制了效率和对未见风格的泛化能力。我们提出了一种轻量级风格条件框架,通过超网络生成的LoRA参数动态调节预训练扩散模型。一个风格参考动作被编码为全局风格嵌入,通过超网络映射到每个去噪步骤的低秩更新。通过监督对比损失结构化风格潜在空间,我们的框架可靠地捕捉了多样的风格属性,提高了对未见风格的泛化能力,并支持基于优化的指导,而无需预定义的风格类别。在HumanML3D和100STYLE数据集上的实验显示,取得了最先进的风格化结果,同时实现了对未见风格的改进风格化。
英文摘要:
Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of motion, commonly referred to as style. Recent approaches have tackled this challenge by attaching a style injection mechanism to a pretrained text-driven diffusion model. Existing stylization methods, however, either require style-specific fine-tuning of existing models or rely on heavy ControlNet-based architectures, limiting efficiency and generalization to unseen styles. We propose a lightweight style conditioning framework that dynamically modulates a pretrained diffusion model through hypernetwork-generated LoRA parameters. A style reference motion is encoded into a global style embedding, which is mapped by a hypernetwork to low-rank updates applied at each denoising step of the diffusion model. By structuring the style latent space with a supervised contrastive loss, our framework reliably captures diverse stylistic attributes, improves generalization to unseen styles, and supports optimization-based guidance without requiring predefined style categories. Experiments on the HumanML3D and 100STYLE datasets show state-of-the-art stylization results, while achieving improved stylization for unseen styles.