发表机构
Peking University; Light Origins; Alibaba(北京大学; 光起源; 阿里巴巴)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CleanMDM提出统一多模态运动清理框架,通过掩码条件生成支持噪声3D运动、稀疏2D/3D关键帧及文本的任意组合,结合LMQD判别器与接触投影优化,在多个数据集上超越现有清理与生成基线。
AI 中文摘要
运动捕捉数据很少能直接使用,因为它们通常表现出缺失片段、抖动、漂移和接触伪影。传统上,损坏的运动由动画师通过从噪声运动中手动识别关键帧、随后进行关键帧校正以及在校正后的关键帧之间插值以重建连贯运动来清理。虽然生成式运动模型的兴起使得自动清理变得可行,但大多数方法作为黑盒去噪器运行,可控性有限,难以保留可靠片段或强制执行特定用户意图。受动画工作流程的启发,我们提出了CleanMDM,一个统一的多模态运动清理框架,将清理问题表述为带有即插即用条件的掩码条件生成。该单一模型支持噪声3D运动、稀疏2D关键帧、稀疏3D关键帧和文本的任意组合。这种设计既支持无需额外用户注释的自动清理,也支持在多模态引导下的可控清理。为了进一步提高运动真实感,我们引入了潜在运动质量判别器(LMQD),以更好地匹配运动学分布并减少滑动、抖动和相互穿透伪影,并且我们应用网格感知接触投影作为测试时优化步骤,以增强接触和物理一致性。跨多个数据集的实验表明,CleanMDM持续优于先前的清理和生成基线,并且低成本条件(文本和2D关键帧)在多模态清理场景中提供了可靠的可控性提升。
英文摘要
Motion capture data is rarely directly usable, as they typically exhibit missing segments, jitter, drift and contact artifacts. Traditionally, corrupted motions are cleaned by animators through the manual identification of keyframes from noisy motion, subsequent keyframe correction, and interpolation between corrected keyframes to reconstruct coherent motion. While the rise of generative motion models has made automatic cleanup feasible, most approaches operate as black box denoisers with limited controllability, making it difficult to preserve reliable segments or enforce specific user intents. Inspired by animation workflows, we present CleanMDM, a unified multimodal motion cleanup framework that formulates cleanup as masked conditional generation with plug-and-play conditions. This single model supports arbitrary combinations of noisy 3D motion, sparse 2D keyframes, sparse 3D keyframes, and text. This design enables both automatic cleanup without additional user annotation and controllable cleanup under multimodal guidance. To further improve motion realism, we incorporate the Latent Motion Quality Discriminator (LMQD) to better match kinematic distributions and reduce skating, jitter, and interpenetration artifacts, and we apply Mesh-Aware Contact Projection as a test-time optimization step to enhance contact and physical consistency. Experiments across multiple datasets demonstrate that CleanMDM consistently outperforms prior cleanup and generation baselines, and that low cost conditions (text and 2D keyframes) provide reliable controllability gains in multimodal cleanup scenarios.