arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniMo:统一人类与动物运动生成

UniMo: Unifying Human and Animal Motion Generation

Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, Yiran Wang, Danning Li, Ian Reid, Richard Hartley

arXiv 2609.12342首次发表:更新:

发表机构

The Australian National University; Adelaide University; Westlake University; The University of Sydney; The Hong Kong University of Science and Technology (Guangzhou); Mohamed bin Zayed University of Artificial Intelligence(澳大利亚国立大学; 阿德莱德大学; 西湖大学; 悉尼大学; 香港科技大学(广州); 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出UniMo统一框架,通过点云表示和动态采样解决骨骼拓扑差异,并构建大规模数据集UniML3D,实现人类与动物运动生成的最先进性能。

AI 中文摘要

3D运动的条件生成因其在机器人、AR/VR、游戏和内容创作中的广泛应用而成为关键研究课题。然而,将文本驱动的人类运动生成的最新进展扩展到动物领域仍面临挑战,原因在于两个核心限制。首先,与标准人体结构不同,动物表现出高度多样化的骨骼拓扑结构,这使得跨物种的统一建模变得困难,并导致逐物种模型效率低下。其次,现有的动物运动数据集在规模和标注质量上存在局限,制约了模型性能。为解决这些挑战,我们提出了UniMo,一个统一的基于点云的运动生成框架,通过将参数化骨骼转换为非参数化表示来绕过拓扑差异,并通过动态采样为活跃关节分配更多点来进一步增强。此外,我们提出了UniML3D,一个涵盖人类和动物类别的大规模运动-语言数据集,包含145,907个运动序列和433,388条描述,比现有动物数据集大102倍以上。我们的方法在UniML3D以及包括HumanML3D、KIT-ML和AnimalML3D在内的三个公开基准上取得了最先进的结果,证明了统一人类-动物运动生成的可行性和有效性。网站:此https URL。

英文摘要

The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending recent advances in text-driven human motion generation to the animal domain remains challenging due to two core limitations. First, animals exhibit highly diverse skeletal topologies, unlike the standard human structure, making unified modeling across species difficult and leading to inefficient per-species models. Second, existing animal motion datasets suffer from limited scale and annotation quality, constraining model performance. To address these challenges, we propose UniMo, a unified point cloud-based motion generation framework that bypasses topological discrepancies by converting parametric skeletons into unparametric representations, further enhanced by dynamic sampling that allocates more points to active joints. Additionally, we present UniML3D, a large-scale motion-language dataset spanning both human and animal categories, containing 145,907 motion sequences and 433,388 captions-over 102x larger than existing animal datasets. Our method achieves state-of-the-art results on UniML3D and three public benchmarks including HumanML3D, KIT-ML, and AnimalML3D, demonstrating the feasibility and effectiveness of unified human-animal motion generation. Website: https://steve-zeyu-zhang.github.io/UniMo.

CommentsAccepted to SIGGRAPH Asia 2026 Posters

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑