arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于3D视觉几何的玻璃表面检测

Glass Surface Detection Grounded in 3D Visual Geometry

Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau

arXiv 2608.26752首次发表:更新:

发表机构

Jiangnan University; University of Science and Technology of China; University of Wurzburg; City University of Hong Kong (Dongguan)(江南大学; 中国科学技术大学; 维尔茨堡大学; 香港城市大学(东莞))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将玻璃表面检测(GSD)基于3D视觉几何的新范式,结合VGGT、FSAM与GeGB,在7个基准上达最优性能,泛化性好且提升玻璃场景重建效果。

AI 中文摘要

玻璃表面检测(GSD)对场景理解与重建至关重要,但因玻璃表面的透明性与反射性仍具挑战性。现有GSD方法通常依赖2D外观线索,在几何模糊场景中可能失效。本文提出范式转变:将GSD基于3D视觉几何,以显式建模玻璃表面的物理存在。我们的方法首先从视觉几何基础Transformer(VGGT)中提取丰富的3D先验,生成感知玻璃的3D表示;随后通过多任务学习,采用新型玻璃检测头,该头包含两个核心模块:识别玻璃特定光谱特征以定位玻璃表面的频率自注意力模块(FSAM),以及选择性将2D特征基于3D几何以分割玻璃表面的几何基础块(GeGB)。大量实验表明,我们的方法在7个标准GSD基准上达到了当前最优性能,对视频/多模态数据泛化性良好,且显著提升了玻璃场景的重建效果。代码可在该https URL获取。

英文摘要

Glass surface detection (GSD) is critical for scene understanding and reconstruction, and yet remains challenging due to the transparency and reflectivity of glass surfaces. Existing GSD methods typically rely on 2D appearance cues, which may fail in geometrically ambiguous scenes. In this paper, we propose a paradigm shift: grounding GSD in 3D visual geometry to explicitly model the physical existence of glass surfaces. Our method first distills rich 3D priors from the visual geometry grounded transformer (VGGT) and generates glass-aware 3D representations. It then exploits multi-tasking learning with a novel glass detection head, consisting of two core modules: a Frequency Self-Attention Module (FSAM) that identifies glass-specific spectral features for glass surface localization, and a Geometry Grounding Block (GeGB) that selectively grounds 2D features in 3D geometry for glass surface segmentation. Extensive experiments demonstrate that our method achieves state-of-the-art performance across seven standard GSD benchmarks, generalizes well to video/multi-modal data, and substantially improves reconstruction in glass scenes. Code is available in https://github.com/YT3DVision/VGGT_GLASS.

Comments9 pages, 10 figures. Accepted by ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑