arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08068cs.CV

PhysTacGen:物理感知的视觉-触觉传感器图像生成

PhysTacGen: Physics-Aware Visual-Tactile Sensor Image Generation

Guo Tang, Yongtao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

PhysTacGen提出了一种结合材料感知描述与几何条件的视觉到光学触觉图像生成框架,通过组触觉策略优化和SDXL ControlNet,在SSVTP数据上提升了结构相似性并验证了生成触觉输入在下游任务中的效用。

中文摘要 AI 辅助

真实的物理交互是具身智能的基石,然而收集配对的视觉-触觉数据成本高昂。视觉到触觉的合成提供了一种有前景的方法来增强此类数据,但学习这种映射因视觉外观与接触相关的材料属性之间的差距以及配对观测中的空间错位而变得复杂。为解决这些挑战,我们提出了PhysTacGen,一个视觉到光学触觉图像生成框架,该框架将材料感知描述与几何条件相结合。首先,我们引入了组触觉策略优化(GTPO),一种强化学习策略,通过任务特定的奖励来优化视觉-语言模型,以生成结构化的材料描述。其次,我们将基于DINOv2的配对筛选与单目相对深度估计相结合,以选择训练配对并提供几何先验。最后,一个SDXL ControlNet基于RGB、相对深度和GTPO生成的文本合成光学触觉图像。在精选的SSVTP数据上的实验表明,与对比基线相比,结构相似性有所提高,而一项盲法用户研究显示对GTPO生成的描述有偏好。生成的触觉输入还提高了基于属性的力系数预测代理的性能。这些结果共同证明了PhysTacGen在光学触觉图像合成方面的有效性及其在评估的下游任务中的实用性。代码将在以下网址提供。

英文摘要

Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual--tactile data remains costly. Visual-to-tactile synthesis offers a promising approach to augmenting such data, but learning this mapping is complicated by the gap between visual appearance and contact-related material properties, as well as spatial misalignment in paired observations. To address these challenges, we present \textbf{PhysTacGen}, a visual-to-optical-tactile image generation framework that integrates material-aware descriptions with geometric conditioning. First, we introduce Group Tactile Policy Optimization (GTPO), a reinforcement learning strategy that refines a vision--language model to generate structured material descriptions using task-specific rewards. Second, we combine DINOv2-based pair curation with monocular relative-depth estimation to select training pairs and provide geometric priors. Finally, an SDXL ControlNet synthesizes optical tactile images conditioned on RGB, relative depth, and GTPO-generated text. Experiments on curated SSVTP data demonstrate improved structural similarity over the compared baselines, while a blinded user study shows a preference for GTPO-generated descriptions. Generated tactile inputs also improve performance on an attribute-derived force-coefficient prediction proxy. Together, these results demonstrate the effectiveness of PhysTacGen for optical tactile image synthesis and its utility in the evaluated downstream task.The code will be available at https://github.com/VDIGPKU/PhysTacGen.

发表机构

  • Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所)
  • VGI Labs Co., Ltd.(VGI Labs 有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑