PAMI:用于文本到人-物体交互生成的部分锚定运动
PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation
查看机构详情
- Tübingen AI Center, University of Tübingen(图宾根大学图宾根人工智能中心)
- Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
PAMI提出部分锚定运动框架,通过身体部位锚点投票表示物体运动并分层细化,实现更忠实的人-物体交互生成,接触召回率提升14.5%。
中文摘要 AI 辅助
文本条件下的全身人-物体交互(HOI)生成需要合成与输入文本匹配的人体运动和物体轨迹,同时保持随时间精确协调。大多数方法将人和物体表示为独立的轨迹,并预测全局的人-物体耦合。然而,隐式学习这种复杂、动态变化的关系往往会导致物体漂移、接触遗漏和穿透。我们提出PAMI,一个用于交互生成的部分锚定运动框架。受经典霍夫变换启发,我们的关键思想是通过让身体部位锚点对物体运动进行投票来定位物体运动:我们将物体运动表示为相对于多个身体部位锚点的运动,并使用PamiVAE学习一个交互潜空间,解码出逐帧权重以聚合这些部位特定的投票。基于这种表示,PAMI以从粗到细的层次生成交互。PamiGen首先在此结构化潜空间中从文本生成粗略的人-物体交互,随后PamiRefiner使用混合表面感知表示递归地解决细粒度接触几何问题,该表示结合了捕捉整体身体部位影响的长距离探针和解决物体表面附近详细接触的短距离传感器。在InterAct上的实验表明,PAMI生成的交互比先前方法更忠实,人相对物体的运动更准确,接触召回率比先前最先进方法高出14.5%。广泛的消融实验验证了部分锚定投票表示和混合表面感知细化两者的贡献。
英文摘要
Text-conditioned full-body human-object interaction (HOI) generation requires synthesizing human motion and object trajectories that match the input text while remaining precisely coordinated over time. Most methods represent the human and object as separate trajectories and predict the global human-object couplings. Learning this complex, dynamically changing relationship implicitly, however, often yields object drift, missed contact, and penetration. We introduce PAMI, a Part-Anchored Motion framework for Interaction generation. Inspired by the classic Hough Transform, our key idea is to localize object motion by letting body-part anchors vote for it: we express object motion relative to multiple body-part anchors and use PamiVAE to learn an interaction latent space, decoding frame-wise weights that aggregate these part-specific votes. Building on this representation, PAMI generates interactions in a coarse-to-fine hierarchy. PamiGen first generates a coarse human-object interaction from text in this structured latent space, and PamiRefiner then recursively resolves fine-grained contact geometry using a hybrid surface-sensing representation, combining long-range probes that capture overall body-part influence with short-range sensors that resolve detailed contacts near the object surface. Experiments on InterAct show that PAMI generates more faithful interactions and more accurate human-relative object motion than previous methods, achieving 14.5% higher contact recall than the previous state of the art. Extensive ablations validate the contributions of both the part-anchored voting representation and hybrid surface-sensing refinement.