发表机构
Hefei University of Technology; HiDream.ai Inc.(合肥工业大学; HiDream.ai公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出RASA框架,通过向DiT注入结构化先验解耦空间映射与运动控制,构建CIM-Bench基准,实验显示其在跨身份角色动画的运动保真度和视觉质量上优于现有最优方法。
AI 中文摘要
跨身份角色动画旨在驱动目标身份,使其从参考图像出发,跟随驱动视频中源角色的运动。核心挑战在于两种能力的固有纠缠:跨身份空间映射(对齐位置、尺度和骨骼比例)与运动控制(优化关节铰接、体积一致性和视图连贯性)。我们提出Reference-Aware Structural Alignment(RASA,参考感知结构对齐)框架,通过向扩散Transformer(DiT)注入结构化先验,将空间映射与运动控制解耦。我们的方法分为两个阶段:第一,空间先验校准器(SPC)将参考身份与驱动姿态融合,生成空间接地的初始噪声潜变量,确保与驱动骨骼的正确定位、缩放和对齐;第二,固有运动引导器(IMG)将与形状无关的SMPL铰接参数编码为语义运动向量,该向量超越外观偏向的2D关键点,注入DiT中间层后,可补充基础姿态条件,实现符合解剖学的铰接和视图感知的体积优化。我们精心策划了经过严格筛选的高质量基准CIM-Bench用于评估。大量实验表明,RASA在运动保真度和视觉质量上显著优于现有最优方法。本研究确立了新范式,证明解耦的空间与运动先验是实现鲁棒角色动画的关键。项目页面:this https URL
英文摘要
Cross-identity character animation aims to drive a target identity from a reference image to follow the motion of a source character from a driving video. The core challenge lies in the inherent entanglement of two capabilities: cross-identity spatial mapping (aligning position, scale, and skeletal proportions) and motion control (refining joint articulation, volumetric consistency, and view coherence). We introduce Reference-Aware Structural Alignment (RASA), a framework that disentangles spatial mapping from motion control by injecting structured priors into a Diffusion Transformer (DiT). Our approach has two stages. First, a Spatial Prior Calibrator (SPC) fuses reference identity with driving pose to generate a spatially grounded initial noise latent, ensuring correct positioning, scaling, and alignment with the driving skeleton. Second, an Inherent Motional Guider (IMG) encodes shape-agnostic SMPL articulation parameters into a semantic motion vector beyond appearance-biased 2D keypoints. Injected into intermediate DiT layers, this vector complements the base pose condition for anatomically consistent articulation and view-aware volumetric refinement. We curate CIM-Bench, a high-quality benchmark with rigorous curation, for evaluation. Extensive experiments show RASA significantly outperforms state-of-the-art methods in motion fidelity and visual quality. Our work establishes a new paradigm showing disentangled spatial and motional priors are key to robust character animation. Project page: https://hidream.ai.github.io/RASA/
CommentsAccepted by ECCV 2026. 15 pages, 7 figures