arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从何处到如何:基于第一人称视频的连续4D交互预测

From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie

arXiv 2609.08636首次发表:更新:

发表机构

School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen); Pengcheng Laboratory; School of Software, Shandong University(哈尔滨工业大学(深圳)计算机科学与技术学院; 鹏城实验室; 山东大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Coherent4D数据集和HIGFlow框架,通过级联的从何处到如何过程,实现基于第一人称视频的连续4D交互位置与全身姿态预测。

AI 中文摘要

第一人称4D交互预测旨在预测未来交互将在3D空间中何处发生以及人体将如何运动以实现这些交互,为辅助机器人和人机交互提供了重要能力。现有方法难以将语义理解转化为精确的连续3D定位,并在姿态预测中平衡运动多样性与结构一致性。更根本的是,这些任务通常被分开建模,导致交互位置与身体运动之间的连续几何和时间对应关系未被充分捕获。为解决这些挑战,我们引入了Coherent4D,一个用于连续4D交互预测的大规模第一人称数据集,包含三个领域的约233K个样本。每个样本将一系列未来3D交互位置与相应的全身姿态配对,这些姿态在时间上对齐并表达在共享坐标系中。我们还提供了连续空间中的评估指标。基于这一表述,我们提出了HIGFlow,一个手部交互引导的残差流框架,将预测建模为级联的从何处到如何的过程。HIGFlow首先通过结合语义锚定与短时视觉动态来预测连续的未来交互位置,然后利用预测的位置序列来调节确定性运动锚点和残差流匹配,以实现多样且结构一致的全身运动预测。在所有三个领域的大量实验表明,在位置和姿态预测方面均持续优于代表性基线,而消融研究验证了所提出组件的贡献。项目页面可从此https URL访问。

英文摘要

Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important capability for assistive robotics and human-computer interaction. Existing methods struggle to translate semantic understanding into precise continuous 3D localization and to balance motion diversity with structural consistency in pose forecasting. More fundamentally, these tasks are often modeled separately, leaving the continuous geometric and temporal correspondence between interaction locations and body motion insufficiently captured. To address these challenges, we introduce Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains. Each sample pairs a sequence of future 3D interaction locations with corresponding full-body poses, aligned in time and expressed in a shared coordinate system. We also provide evaluation metrics in continuous space. Building on this formulation, we propose HIGFlow, a Hand Interaction Guided Residual Flow framework that models forecasting as a cascaded where-to-how process. HIGFlow first forecasts continuous future interaction locations by combining semantic grounding with short-horizon visual dynamics, and then uses the predicted location sequence to condition a deterministic motion anchor and residual Flow Matching for diverse yet structurally consistent full-body motion forecasting. Extensive experiments across all three domains demonstrate consistent improvements over representative baselines on both location and pose forecasting, while ablations validate the contributions of the proposed components. The project page is available at https://corrineqiu.github.io/from-where-to-how/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑