AI 中文总结
研究在视觉不可用场景下从稀疏触摸重建可变形物体完整网格的问题,提出用排列不变交叉注意力架构的估计器,相比其他方法大幅降低重建误差,还能利用其不确定性指导下次触摸位置,进一步降错。
AI 中文摘要
在视觉不可用时估计可变形物体的完整形状极具挑战性,如在黑暗中、不透明袋子里等。触摸是这些场景中的自然传感器,但触摸是稀疏且局部的。我们提出一种单一的与拓扑无关的估计器,仅通过几次触摸且无视觉信息就能重建可变形物体的完整网格,采用一种处理一维绳索、二维布料和三维体积软体的排列不变交叉注意力架构。学习到的估计器相对于非学习的几何网格完成和高斯过程曲面基线,将重建误差降低了约三分之二,且优于更简单的全局池集编码器。此外,估计器的深度集成不确定性可用于学习下一次触摸位置,在稀疏预算下进一步降低误差,优于随机触摸和高斯过程主动基线。当有视觉信息时,触摸位置影响不大,凸显了我们所研究的无视觉设置的重要性。
英文摘要
Estimating the full shape of a deformable object is especially challenging when vision is unavailable: in the dark, inside an opaque bag, behind the manipulating hand, or under heavy self-occlusion. Touch is the natural sensor in these settings, but touches are sparse and local. We present a single topology-agnostic estimator that reconstructs the full mesh of a deformable object from only a few touches and no vision, using one permutation-invariant cross-attention architecture that handles a 1D rope, a 2D cloth, and a 3D volumetric soft body. The learned estimator reduces reconstruction error by roughly two-thirds relative to non-learned geometric mesh completion and a Gaussian-process surface baseline, and it outperforms a simpler global-pool set encoder, with the gap growing as more touches are observed. We then show that the estimator's deep-ensemble uncertainty can be used to learn where to touch next, which lowers error further and beats both random touching and a Gaussian-process active baseline at sparse budgets. This gain is modest on average but grows with self-occlusion and on the error tail. When vision is also available, where to touch barely matters, motivating the vision-free setting we study.
CommentsAccepted to the RSS 2026 DAROMA Workshop