arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoSAT:基于空间音频和文本描述的人体运动生成

MoSAT: Human Motion Generation from Spatial Audio and Textual Description

Shuyang Xu, Zhiyang Dou, Yiduo Hao, Zekun Li, Liang Pan, Jingbo Wang, Cheng Lin, Yuan Liu, Wenping Wang, Mingmin Zhao, Taku Komura

arXiv 2609.23797首次发表:更新:

发表机构

The University of Hong Kong; Massachusetts Institute of Technology; University of Pennsylvania; Brown University; Shanghai AI Lab; Macau University of Science and Technology; Hong Kong University of Science and Technology; Texas A&M University(香港大学; 麻省理工学院; 宾夕法尼亚大学; 布朗大学; 上海人工智能实验室; 澳门科技大学; 香港科技大学; 德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MoSAT,一种基于空间音频和文本描述联合生成人体运动的潜在流匹配框架,并引入STAM数据集,通过分层交叉注意力实现最先进的运动生成性能。

AI 中文摘要

人体运动既受外部声学事件影响,也受行为意图驱动:空间音频传递环境线索,引发或引导响应,而文本则指定期望的动作及其执行方式。本文研究了一项新颖任务——联合基于空间音频和自然语言的人体运动合成,这一问题在以往研究中很大程度上被忽视。为支持该任务,我们引入了STAM数据集,包含与空间音频和详细文本标注配对的动作序列,其丰富的词汇支持对人体运动进行精确且细致的描述。我们进一步提出了MoSAT,一种潜在流匹配框架,用于全身运动生成,该框架通过分层交叉注意力,在生成运动前联合基于自然语言意图和定向空间音频线索进行条件控制。这种分层设计增强了时间连贯且语义对齐的运动序列。我们还开发了三模态评估器,用于对这一新颖任务进行全面评估。大量实验表明,MoSAT通过利用空间音频固有的运动塑造特性以及文本语义,实现了最先进的性能,能够在各种场景中生成精确且多样的运动。

英文摘要

Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial audio and natural language, a problem that has been largely overlooked in previous research. To support this task, We introduce STAM, a dataset of motion sequences paired with spatial audio and detailed textual annotations whose rich vocabulary affords precise and nuanced specification of human motions. We further introduce MoSAT, a latent flow-matching framework for full-body motion generation jointly conditioned on natural-language intent and directional spatial-audio cues through hierarchical cross-attention before generating motion. Such a hierarchical design enhances temporally coherent and semantically aligned motion sequences. We also develop tri-modal evaluators for comprehensive evaluation on this novel task. Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio's intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑