arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29475cs.CV

从极端视觉稀疏性中洞察:基于单个随机视觉斑块的表面理解

Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch

Sindhuja Penchala, Sudip Mittal, Noorbakhsh Amiri Golilarz

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出SSUF框架,适配四种预训练架构,在仅保留10%可见区域的Touch-and-Go数据集上,实现极端视觉稀疏下的表面重建与材质分类,各模型表现不同且均达实时推理。

中文摘要 AI 辅助

从不完整的视觉观测中识别表面材质,仍是机器人感知与环境理解领域的一个挑战性问题。本文提出稀疏表面理解框架(Sparse Surface Understanding Framework, SSUF),这是一个统一的双任务学习框架,它适配了四种预训练架构——卷积自编码器(Convolutional Autoencoder, ConvAE)、视觉Transformer(Vision Transformer, ViT)、Swin Transformer和掩码自编码器(Masked Autoencoder, MAE),以同时实现表面重建与材质分类。实验在Touch-and-Go数据集上开展,采用稀疏观测协议,仅保留原始图像的10%可见区域,其余区域均被掩码。为实现公平对比,面向重建的模型扩展了分类头,而面向分类的模型则添加了重建解码器。所得架构通过重建质量、分类性能、模型复杂度与推理效率指标进行评估。实验结果显示各模型具有不同优势:Swin Transformer取得最优分类性能,准确率达89.21%、F1分数为0.8922、ROC-AUC为0.9813;MAE在评估模型中获得最高重建分数,PSNR为16.06 dB、SSIM为0.4501;ViT则在重建与分类性能间实现最佳整体平衡。此外,所有模型均实现实时推理,单张图像推理耗时不足5 ms。总体而言,结果表明预训练架构可支持极端视觉稀疏条件下的材质识别,但精准图像重建仍具挑战性。

英文摘要

Surface material recognition from incomplete visual observations remains a challenging problem in robotic perception and environmental understanding. This paper discusses Sparse Surface Understanding Framework (SSUF), a unified dual-task learning framework that adapts four pretrained architectures-Convolutional Autoencoder (ConvAE), Vision Transformer (ViT), Swin Transformer, and Masked Autoencoder (MAE) for si-multaneous surface reconstruction and material classification. Experiments were conducted on the Touch-and-Go dataset using a sparse observation protocol in which only 10% of the original image remained visible while the remaining regions were masked. To enable a fair comparison, reconstruction-oriented models were extended with classification heads, whereas classification- oriented models were augmented with reconstruction decoders. The resulting architectures were assessed using reconstruction quality, classification performance, model complexity, and in-ference efficiency metrics. Experimental results revealed distinct strengths across the models. Swin Transformer achieved the best classification performance with an accuracy of 89.21%, an F1-score of 0.8922, and a ROC-AUC of 0.9813. In contrast, MAE produced the highest reconstruction scores among evaluated models, with a PSNR of 16.06 dB and an SSIM of 0.4501, while ViT provided the best overall balance between reconstruction and classification performance. Furthermore, all models achieved real-time inference, requiring less than 5 ms per image. Over-all, the results show that pretrained architectures can support material recognition under severe visual sparsity, while accurate image reconstruction remains challenging.

发表机构

  • The University of Alabama(阿拉巴马大学)

机构由 AI 辅助整理,请以论文原文为准。

↑