寄生式协同去噪:在冻结的视频扩散模型中解锁三维人体运动生成
Parasitic Co-Denoising: Unlocking 3D Human Motion Generation in a Frozen Video Diffusion Model
浏览论文内容
中文总结 AI 辅助
本文提出寄生式协同去噪范式,通过冻结视频扩散模型中间特征解码三维人体运动,以极少参数实现文本-运动对齐并同步生成视频与运动。
中文摘要 AI 辅助
尽管从未在显式三维运动数据上进行监督,大规模文本到视频扩散模型在其生成的视频中能够合成逼真的人体运动。我们探究是否可以将这种隐式知识转化为显式的三维运动生成,而无需训练单独的运动模型。对冻结的Wan2.1进行探测揭示,其整个去噪调度过程中的中间状态中存在可恢复的运动信号,而不仅限于干净输出。受此启发,我们引入了寄生式协同去噪,这是一种从宿主模型沿其去噪调度解码运动而非由独立生成器产生运动的范式。我们将其实例化为寄生运动解码器(PMD),一种高效的流匹配解码器,它共享宿主的噪声调度,并通过σ自适应多层融合读取其中间特征,保持宿主不变。PMD从宿主而非运动数据中获取覆盖范围,以极少的可训练参数在文本-运动对齐上领先于专用运动生成器,同时以运动-only基线无法匹敌的单次生成配对视频和运动。
英文摘要
Despite never being supervised on explicit 3D motion, large-scale text-to-video diffusion models synthesize realistic human motion in their generated videos. We ask whether this implicit knowledge can be turned into explicit 3D motion generation, without training a separate motion model. Probing a frozen Wan2.1 reveals that a recoverable motion signal is present in its intermediate states across the entire denoising schedule, not confined to the clean output. Motivated by this, we introduce parasitic co-denoising, a paradigm in which motion is decoded from the host model along its denoising schedule rather than produced by an independent generator. We instantiate it as the Parasitic Motion Decoder (PMD), an efficient flow-matching decoder that shares the host's noise schedule and reads its intermediate features through a $σ$-adaptive multi-layer fusion, leaving the host unmodified. Drawing its coverage from the host rather than from motion data, PMD leads dedicated motion generators on text-motion alignment at a small fraction of their trainable parameters, while producing paired video and motion in a single pass that motion-only baselines cannot match.
发表机构
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。