发表机构
Roblox; UMass Amherst; TU Crete(罗布乐思; 马萨诸塞大学阿默斯特分校; 克里特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BLARM方法,无需骨骼等标注,通过融合刚性运动原语实现视频驱动的3D网格动画,经多任务训练可生成准确稳定且结构可解释的动画。
AI 中文摘要
我们提出BLARM,一种用于视频驱动3D网格动画的前馈方法。给定单目视频和静态物体网格,BLARM会生成时间连贯的动画网格,其运动与视频一致。该方法不依赖显式绑定或直接回归高维顶点运动,而是用一组紧凑的、学习得到的时变刚性运动分量,以及时不变的顶点-分量蒙皮权重来表示动画,从而形成低维变形空间,无需骨骼、 cage( cage模型)、蒙皮权重或绑定标注。其架构通过分解时空注意力将几何衍生的变形潜态条件于视频特征,再解码由预测蒙皮权重融合的刚性变换。经轨迹重构、熵正则化和运动感知对比学习训练后,BLARM能生成准确且时间稳定的动画,同时从单目视频中恢复紧凑、可解释的运动结构。
英文摘要
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.