发表机构
University of Bonn; Lamarr Institute for Machine Learning and Artificial Intelligence(波恩大学; 拉马尔机器学习与人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对复杂场景下多人运动预测中物体信息与社会交互整合的挑战,提出OCSD模型,在HiK和HOI-M3基准上实现最优路径误差降低,生成更逼真长期预测。
AI 中文摘要
在复杂场景中准确预测人员的运动需要对整个环境的过去和当前状态进行推理,在此背景下,将物体信息和社会交互有效整合到统一框架中仍极具挑战性。为解决该问题,我们提出了物体条件社交扩散模型(Object-Conditioned Social Diffusion, OCSD),这是一种将运动历史、多人交互和物体线索整合到单一框架中的条件扩散模型。OCSD采用物体条件机制在每个时间步调节去噪过程,实现细粒度的人-物推理,同时采用社交编码器对场景中所有人类之间的交互进行建模。因此,该模型可自然处理不同的群体规模、复杂的社会交互,并支持采样多个合理的未来轨迹。大量实验表明,OCSD在厨房人类(Humans in Kitchens, HiK)和HOI-M3基准测试中取得了最先进的结果;与现有工作相比,其在HiK上将2秒路径误差降低了121.5毫米(31.3%),在HOI-M3上降低了130.5毫米(33.2%),并生成更逼真的长期预测。
英文摘要
Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.
CommentsAccepted to GCPR 2026