AI 中文总结
研究旨在解决现有触觉生成方法依赖特定传感器数据集且泛化能力不足的问题。提出VQ-Touch框架,含DM-VQGAN提取特征及离散扩散解码器支持多模态生成,经少样本混合训练增强泛化能力,实验显示其在多任务中超越现有方法。
AI 中文摘要
触觉图像生成通过合成高保真触觉数据,显著降低了对昂贵且易磨损传感器的依赖,为机器人感知和人机交互系统中的触觉信息获取提供了有效解决方案。然而,现有方法依赖特定传感器的大规模多样数据集,缺乏高效数据利用和强大泛化能力,在视觉受限环境中表现不佳。为解决此问题,我们引入了VQ-Touch,一个支持跨传感器和多场景应用的触觉生成框架。具体而言,为了从数据中高效提取复杂变形和纹理特征,我们提出了DM-VQGAN,一种有效的触觉表示学习器。此外,我们引入了具有统一条件接口的离散扩散解码器,支持图像和标签等多模态生成任务,并通过少样本混合训练增强模型的泛化能力,并实现与当前主流传感器及其变体的兼容性。实验表明,VQ-Touch在多个任务中超越了现有方法。
英文摘要
Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic perception and human-machine interaction systems. However, existing methods depend on large-scale, diverse datasets from specific sensors and lack efficient data utilization and robust generalization capabilities, struggling in vision-limited environments. To address this, we introduce VQ-Touch, a tactile generation framework that supports both cross-sensor and multi-scenario applications. Specifically, to efficiently extract complex deformation and texture features from the data, we propose DM-VQGAN, an effective tactile representation learner. Furthermore, we introduce a discrete diffusion decoder with a unified conditioning interface, supporting multimodal generation tasks such as images and labels, and enhances the model's generalization capability through few-shot mixed training, thus achieving compatibility with current mainstream sensors and their variants. Experiments show that VQ-Touch surpasses state-of-the-art methods in multiple tasks.
Comments6 pages, 5 figures