ReMoMask-2:潜在检索增强的掩码运动生成
ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation
浏览论文内容
中文总结 AI 辅助
提出ReMoMask-2,通过潜在空间检索和结构感知融合,实现高效准确的文本到运动生成,在多个基准上达到最优性能。
中文摘要 AI 辅助
文本到运动(T2M)生成将自然语言映射为人体关节运动,助力游戏、虚拟现实和机器人技术。检索增强的文本到运动(RAG-T2M)通过基于检索到的运动-文本对进行条件生成,提升了对复杂描述的处理能力。然而,现有的RAG-T2M模型面临两个挑战:粗粒度的检索和融合机制忽视了人体运动的层次化、时空拓扑结构;同时,由于检索到的证据存在于与生成器潜在空间分离的语义空间中,存在表示差距。为解决第一个挑战,我们提出了ReMoMask,一个结构感知的RAG框架,结合了层次化双向动量(HBM)对比学习以对齐全局和部件级特征与文本;语义时空注意力(SSTA)用于拓扑感知融合;以及拓扑结构掩码(TSM)通过自适应掩码强制进行鲁棒的部件级基础。为解决第二个挑战,我们引入了ReMoMask-2,它直接在生成器的预量化潜在空间中重建检索数据库,并通过蒸馏的轻量级投影器对齐文本查询,使生成器能够直接消费检索到的运动的语义内容。在HumanML3D、KIT-ML和SnapMoGen上的大量实验表明,我们的检索器达到了最先进的准确性,而ReMoMask-2在KIT-ML和SnapMoGen上取得了最低的FID;值得注意的是,其单阶段掩码变换器超越了ReMoMask的完整两阶段流水线,并提供了最快的推理速度。
英文摘要
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotics. Retrieval-Augmented Text-to-Motion (RAG-T2M) improves generation on complex descriptions by conditioning on retrieved motion-text pairs. However, existing RAG-T2M models face two challenges: coarse-grained retrieval and fusion mechanisms overlook the hierarchical, spatial-temporal topology of human motion, and a representation gap exists because retrieved evidence resides in a semantic space separate from the generator's latents. To address the first, we present ReMoMask, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum (HBM) contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking. To address the second, we introduce ReMoMask-2, which rebuilds the retrieval database directly within the generator's pre-quantization latent space and aligns text queries via a distilled lightweight projector, allowing the generator to directly consume the retrieved motion's semantic content. Extensive experiments on HumanML3D, KIT-ML, and SnapMoGen demonstrate our retriever achieves state-of-the-art accuracy, while ReMoMask-2 attains the lowest FID on KIT-ML and SnapMoGen; notably, its single mask-transformer stage surpasses ReMoMask's full two-stage pipeline and delivers the fastest inference.
发表机构
- University of Sydney(悉尼大学)
- School of Computer Science, Peking University(北京大学计算机学院)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。