arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HOIMask:面向用于人机交互生成的生成式掩码建模

HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation

Yihong Ji, Jinsong Zhang, He Hu, Hongbo Xu

arXiv 2608.15141首次发表:更新:

发表机构

College of Computer Science and Software Engineering, Shenzhen University; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Fuzhou University(深圳大学计算机与软件学院; 广东省人工智能与数字经济实验室(深圳); 福州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HOIMask作为首个用于离散空间HOI运动建模的生成式掩码框架,通过HOI VQ编码、Transformer架构及接触感知重构引导,生成的HOI运动优于现有扩散方法。

AI 中文摘要

基于扩散的方法在人机交互(HOI)生成领域占据主导地位,因为它们能借助关键的接触融合或信号引导扩散过程。然而,这类方法常因迭代去噪过程中的误差累积,产生大量伪影且交互质量不稳定。本研究提出HOIMask,首个用于在离散空间中建模HOI运动的生成式掩码框架。HOIMask首先通过HOI向量量化(VQ)将运动序列和接触感知信号编码为离散的二维人体与物体token图,保留了超越传统一维表示的细粒度时空结构。在此基础上,采用生成式掩码建模框架联合捕捉人机交互动态,利用专为建模复杂时空与交互依赖关系设计的Transformer架构。为生成更连贯、物理上更合理的运动,我们进一步在推理阶段引入离散空间中新颖的接触感知重构引导,该引导融合接触信号以优化HOI token,迫使生成的运动具备更高的时空一致性。凭借精心设计的运动交互token、专用架构与引导策略,HOIMask在性能上优于当前最优的基于扩散的方法,能生成更逼真、语义对齐的HOI运动。更多结果可参考此https URL。

英文摘要

Diffusion-based methods have dominated the HOI generation, as they enable critical contact fusions or signals to guide the diffusion process. However, they often result in high artifacts and unstable interaction quality due to error accumulation during iterative denoising. In this work, we propose HOIMask, the first generative masked framework for modeling HOI motion in discrete space. HOIMask first encodes both motion sequences and contact-aware signals into discrete 2D human and object token maps via HOI Vector Quantization (VQ), preserving fine-grained spatial-temporal structure beyond conventional 1D representations. On this basis, a generative masked modeling framework is employed to jointly capture human-object interaction dynamics, leveraging a transformer architecture designed to model complex spatial-temporal and interaction dependencies. To generate more coherent and physically plausible motions, we further introduce a novel contact-aware reconstruction guidance in discrete space during inference, which fuses contact signals to optimize HOI tokens that forces the generated motion with higher spatio-temporal consistency. With craftily designed motion interaction tokens, dedicated architecture and guidance strategy, HOIMask outperforms state-of-the-art diffusion-based methods, generating more realistic and semantically aligned HOI motions. Please refer to https://jyhflash.github.io/HOIMask/ for more results.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑