arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38578cs.CVcs.AI

可学习展平实现到多样骨骼的运动重定向

Retargeting Motions to Diverse Skeletons via Learnable Flattening

Kia-Jüng Yang, Fabian H. Sinz, Paweł A. Pierzchlewicz

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种Transformer自编码器,通过可学习展平骨骼图实现跨拓扑运动重定向,在零样本下将全局关节误差降低43-47%,并获用户研究最高评价。

中文摘要 AI 辅助

跨结构运动重定向旨在不同骨骼拓扑之间转移运动。尽管近期有所进展,现有最先进模型在零样本设置(即训练时未见的不同拓扑骨骼)下仍难以保证可靠性,且近期基于Transformer的尝试未能超越专门的几何方法。我们通过一种Transformer自编码器弥合了这一差距,该自编码器学习一个拓扑不变且平移不变的潜在空间。我们的核心贡献是对骨骼图的可学习展平,它同时捕获局部依赖和全局结构。与标准Transformer架构(将位置信息添加到令牌内容中)不同,我们将基于图的位置编码以乘法方式集成,这一设计选择直接源于我们的展平公式。所得模型在单一统一架构内处理多样骨骼拓扑,并以完全无监督方式训练,无需成对重定向数据。消融研究表明,图编码、乘法公式和Transformer主干对性能至关重要。在零样本评估中,我们的方法将全局关节位置误差比当前基准降低了43-47%。一项包含专家动画师的用户研究(n=37)进一步将我们的方法在运动对齐和物理合理性方面评为最高(p<0.05)。这些结果表明,我们的模型设计是使Transformer架构对运动重定向有效的关键,超越了现有方法。

英文摘要

Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by $43-47\%$ over current benchmarks. A user study ($n = 37$), including expert animators, further ranks our approach highest in motion alignment and physical plausibility ($p < 0.05$). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.

发表机构

  • Institute of Computer Science, University of Göttingen(哥廷根大学计算机科学研究所)
  • Campus Institute Data Science, University Göttingen(哥廷根大学数据科学校园研究所)
  • Pantomim P.S.A

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑