发表机构
University of Chinese Academy of Sciences; Tencent HunYuan(中国科学院大学; 腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对HOI视频生成中现有方法依赖显式运动控制的问题,提出AgentHOI,通过多智能体推理和隐式文本-运动对齐策略,实现文本驱动的HOI视频生成,提升了交互自然性等,改进了复杂场景下的表现。
AI 中文摘要
视频扩散模型的进展激发了对人类-物体交互(HOI)视频生成的兴趣,其需要对交互逻辑进行细粒度控制。现有HOI方法严重依赖显式运动控制,限制了跨不同物体和交互的可扩展性和泛化性。本研究提出AgentHOI,一种基于文本驱动的HOI视频生成方法,遵循生成前思考框架,通过多智能体对感知、交互和运动规划的推理弥合高层文本意图与物理执行之间的差距。基于生成的交互计划,进一步加强文本驱动的运动理解。引入隐式文本-运动对齐策略,将文本到运动的先验知识融入视频扩散模型,在推理时无需显式运动输入即可实现强大的HOI合成。实验表明,AgentHOI在具有挑战性的以物体为中心的场景(如穿戴和骑行)中显著提高了交互自然性、物体外观保留率以及对复杂文本指令的遵循程度。
英文摘要
Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control over interaction logic beyond single-subject animation. However, existing HOI methods rely heavily on explicit motion control, limiting scalability and generalization across diverse objects and interactions. In this study, we propose AgentHOI, a text-driven HOI video generation following a thinking-before-generation framework that bridges the gap between high-level textual intent and physical execution through multi-agent reasoning over perception, interaction, and motion planning. Building upon the generated interaction plans, we further strengthen text-driven motion understanding. We introduce an implicit text-motion alignment strategy that distills text-to-motion priors into the video diffusion model, enabling robust HOI synthesis without explicit motion inputs at inference. Experiments show that AgentHOI significantly improves interaction naturalness, object appearance preservation, and adherence to complex textual instructions across challenging object-centric scenarios such as wearing and riding. The code is available at https://github.com/bone-11/agenthoi.
CommentsUnder review. The code is available at https://github.com/bone-11/agenthoi