发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对动画、AR/VR等场景中多人交互的文本-运动表示学习问题,提出多模态交互式运动编码器MIME,通过特定方法捕捉结构并训练,在文本-运动检索任务中表现出色,还能跨数据集支持下游运动生成。
AI 中文摘要
文本-运动表示学习发展迅速,在动画、AR/VR和具身AI的多人交互方面兴趣日增。这些场景需要能将语言与个体演员动态及演员间关系对齐的表示。我们引入多模态交互式运动编码器(MIME),据我们所知,它是首个专为两人交互式运动设计的专用多模态编码器。MIME通过基于流的协同注意力捕捉个体和共享结构,并采用基于课程的对比训练。在Inter-X文本-运动检索中,MIME在不同图库规模下始终优于早期和晚期融合基线,在2000样本图库中,文本到运动的R@1相对提高了12.8%。我们还在未见过的InterHuman数据集上,将MIME作为TIMotion和InterMask中的固定辅助先验进行评估。MIME在TIMotion中提高了语义对齐指标,同时保持了可比的FID。这些结果表明,交互感知多模态编码改善了多人运动检索,并能跨数据集转移以支持下游运动生成。
英文摘要
Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.
CommentsUnder review at WACV 2027