arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用物体几何先验从视觉学习预测接触力分布

Learning to Predict Contact Force Distributions from Vision Leveraging Object Geometry Priors

Ryo Hanai, Yukiyasu Domaea, Ixchel G. Ramirez-Alpizar, Abdullah Mustafa, Floris Erich, Tetsuya Ogata

arXiv 2608.00464首次发表:更新:

AI 中文总结

本文提出融入物体几何先验的模型,从单张RGB图像预测三维接触力分布,经模拟与真实环境评估,该方法提升了力预测精度与下游任务性能,且可从模拟泛化到真实场景。

AI 中文摘要

人类基于视觉和先验经验可做出粗略物理预测并调整操控策略,本文旨在赋予机器人类似能力。为收集视觉与力的配对数据,我们采用机器人领域常用的刚体模拟器;与输出含噪点力的模拟器不同,人类即便在陌生情境也能做出一致预测,据此我们假设:预测平滑力分布而非原始点力,可同时提升力预测本身及下游任务性能。为验证该假设,我们构建了一个模型,能从堆叠日常物体的单张RGB图像预测三维力分布;目标分布通过对模拟器得到的点力应用统计平滑生成,且我们在平滑过程中融入物体几何,以解释接触状态的变化并实现更一致的视觉预测。我们在模拟环境和真实环境中开展了大量评估,结果显示:我们的方法提升了预测精度,通过平滑增强了下游任务性能,且进一步得益于几何引导的平滑;值得注意的是,该模型仅在模拟环境中训练,却能有效泛化到真实世界场景。

英文摘要

Based on vision and prior experience, humans can make rough physical predictions and adjust their manipulation strategies. This paper aims to endow robots with a similar ability. To collect paired data of vision and forces, we use a rigid-body simulator commonly adopted in robotics. However, unlike simulators that output noisy point forces, humans are able to make consistent predictions even in unfamiliar situations. Based on this observation, we hypothesize that predicting smooth force distributions rather than raw point forces can improve both force prediction itself and downstream task performance. To validate this hypothesis, we construct a model that predicts three-dimensional force distributions from a single RGB image of piled daily objects. The target distribution is generated by applying statistical smoothing to point forces obtained from the simulator. Moreover, by incorporating object geometry into the smoothing process, we aim to account for variations in contact states and achieve more consistent vision-based predictions. We conduct extensive evaluations in both simulation and real environments. Results show that our approach improves prediction accuracy, enhances downstream task performance through smoothing, and further benefits from geometry-guided smoothing. Remarkably, the trained model generalizes effectively to real-world scenes despite being trained solely in simulation.

CommentsPublished in Advanced Robotics, 2026

DOI:10.1080/01691864.2026.2693560

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑