发表机构
LY Corporation(LY Corporation)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有运动生成受限于低维潜在空间的问题,提出表示感知流匹配框架MSFlow,直接在连续运动空间生成,并通过RA-MMDiT及投影采样实现最先进性能与零样本控制。
AI 中文摘要
扩散模型和流模型的最新进展显著提升了文本驱动的人体运动生成质量。然而,大多数方法在低维、时间下采样的潜在空间中生成,这些空间主要为了重建而学习,这一瓶颈限制了生成质量,并妨碍了对单个帧和关节的直接操作。我们提出了MotionSpaceFlow(MSFlow),一种表示感知的流匹配框架,它直接在连续运动空间中预测干净的运动,无需学习编码器或解码器。为了考虑直接运动表示的各向异性结构,我们提出了表示感知的噪声缩放,并展示了初始高斯源尺度如何控制中间概率路径边缘的协方差。我们进一步引入了表示感知的多模态扩散Transformer(RA-MMDiT),它通过联合注意力同时更新令牌级语言和全分辨率运动特征,同时使时间信息流适应运动表示:对由帧间变化定义的增量特征使用因果注意力,对绝对关节坐标等全局特征使用双向注意力。在不同的数据集和运动表示上,MSFlow实现了最先进的文本到运动生成性能。其全局表示变体还通过投影采样实现了对任意关节或帧的零样本、推理时控制,无需控制条件训练,以精确约束满足提供领先的运动质量。
英文摘要
Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of individual frames and joints. We introduce MotionSpaceFlow (MSFlow), a representation-aware flow-matching framework that predicts clean motion directly in continuous motion space without a learned encoder or decoder. To account for the anisotropic structure of direct motion representations, we propose representation-aware noise scaling and show how the initial Gaussian source scale governs the covariance of intermediate probability-path marginals. We further introduce a Representation-Aware Multimodal Diffusion Transformer (RA-MMDiT), which jointly updates token-level language and full-resolution motion features through joint attention while adapting temporal information flow to the motion representation: causal attention for incremental features defined by frame-to-frame changes, and bidirectional attention for global features such as absolute joint coordinates. Across different datasets and motion representations, MSFlow achieves state-of-the-art text-to-motion performance. Its global representation variant additionally enables zero-shot, inference-time control over any joint or frame through projection sampling without control-conditioned training, delivering leading motion quality with exact constraint satisfaction.
CommentsProject website: https://yu1ut.com/MSFlow-HP/