arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAD-HOI:用于从文本生成关节手-物体交互的掩码自回归扩散模型

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

Ananya Bal, Kartik Sharma, Ethan Lai, Samyak Tiwari, Liza Dahiya, Chaitanya Chawla, Laszlo A. Jeni

arXiv 2608.10162首次发表:更新:

AI 中文总结

MAD-HOI是一种结合掩码自回归与扩散的HOI生成模型,可生成多样且符合物理规律的手-物体交互序列,支持可变长度生成、运动补全填充等功能,在ARCTIC和GRAB数据集上优于开源基线。

AI 中文摘要

基于文本生成手-物体交互(HOI)序列的方法主要聚焦于生成平滑、符合物理规律的轨迹。而真正实用的方法还应支持可变长度生成、复合运动序列、运动补全与填充,以及可靠的终止功能,同时不损害物理合理性。标准HOI生成扩散模型通常仅针对原子运动的文本到运动生成进行训练,且需预先指定运动长度。自回归(AR)方法提供了更强的序列级灵活性,但通常依赖离散运动编码,可能丢失接触敏感的运动细节。为解决这些关键局限,我们提出了用于HOI生成的掩码自回归扩散模型(MAD-HOI)。我们的方法首先在连续潜在空间中编码手与物体的运动,同时保持其解耦以维持流级控制;随后通过掩码自回归Transformer预测上下文特征,以调节流匹配头。MAD-HOI可生成原子与复合关节序列的运动,支持条件运动补全与填充,还能通过单一训练目标预测运动结束(EOM)。我们对这些能力进行了全面评估,并在ARCTIC和GRAB数据集上对方法进行了基准测试。实验表明,与其他开源基线方法相比,我们的方法能生成更多样、更符合物理规律的交互。

英文摘要

Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A truly utilitarian method should additionally support variable-length generation, composite motion sequences, motion completion and infilling, and reliable termination without compromising physical plausibility. Standard diffusion models for HOI generation are typically trained only for text-to-motion generation on atomic motions and require the motion length to be specified a-priori. Autoregressive (AR) methods provide greater sequence-level flexibility, but commonly depend on discrete motion codes, which can lose contact-sensitive motion detail. To address these key limitations, we present a model performing Masked Autoregression with Diffusion for HOI generation (MAD-HOI). Our method starts by encoding hand and object motions in a continuous latent space while keeping them disentangled to maintain stream-wise control. This is followed by a masked autoregressive transformer to predict context features that condition a flow-matching head. MAD-HOI is capable of motion generation for atomic and composite articulated sequences, conditioned motion completion and infilling, as well as EOM (End of Motion) prediction from a single training objective. We provide comprehensive evaluations for these capabilities and benchmark our method on the ARCTIC and GRAB datasets. Our experiments demonstrate that our method generates more diverse and physically plausible interactions compared to other open-sourced baseline methods.

Comments17 pages, 9 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑