发表机构
School of Artificial Intelligence, Shandong University; Institute of Automation, Chinese Academy of Sciences; School of Mechanical Engineering, Tianjin University(山东大学人工智能学院; 中国科学院自动化研究所; 天津大学机械工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出仅用RGB的EgoPhys框架,结合CASA与TMTR,从第一人称视角视频估计关节对象操作的峰值接触力与机械功,在Hoi!数据集上取得了更优的预测性能。
AI 中文摘要
对关节对象进行基于物理的操作,需要理解接触过程中遇到的最大力以及部件移动时所做的功。峰值接触力和机械功可量化这两个互补方面,但从第一人称视角视频中估计它们颇具挑战性,因为物理交互线索具有局部性和间接性。此外,峰值力与短暂的接触事件相关,而机械功则取决于整个接触持续时间内的力-运动耦合。为应对这些挑战,我们提出EgoPhys,这是一个仅使用RGB的框架,包含接触感知空间聚合(Contact-Aware Spatial Aggregation,CASA)和目标特定多专家时间路由(Target-Specific Multi-Expert Temporal Routing,TMTR)。CASA整合外观和几何特征以强调与交互相关的线索,而TMTR则利用专门的时间专家对语义、事件和运动线索进行建模,并分别对其进行路由以用于力和功的预测。在Hoi!数据集的测试划分上,EgoPhys大幅提升了峰值力和机械功的预测性能,分别达到了5.205±0.584 N和0.894±0.081 J的平均绝对误差(MAE)。
英文摘要
Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of \(5.205 \pm 0.584\) $N$ and $0.894 \pm 0.081$ $J$, respectively.