发表机构
School of Software, Northeastern University, Shenyang, China(软件学院,东北大学,沈阳,中国)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对透明物体给深度估计和3D重建带来的挑战,提出几何引导预处理框架GHOST,利用视觉基础模型,经多步骤处理分离属性、恢复法线先验并整合多模态线索,合成不透明RGB图像,显著提升相关模型对透明物体的处理精度。
AI 中文摘要
透明物体因其违反朗伯假设,给深度估计和3D重建带来了根本挑战,导致下游任务中严重的几何退化。为解决此问题,我们提出了一种新颖的几何引导预处理框架GHOST,它利用视觉基础模型将透明区域转换为不透明、结构一致的表示,且无需下游模型重新训练。具体而言,我们的流程利用TransDINO和TransDecomp来分离掩码和透明度物理属性,同时DAF-Net恢复表面法线先验以编码几何曲率。随后,GeoSemTransNet整合这些多模态线索以合成保留透明物体3D结构的纹理丰富的不透明RGB图像。大量实验表明,我们的方法通过恢复基本的光度线索,显著提高了最先进的深度估计和重建模型对透明物体的准确性。
英文摘要
Transparent objects pose a fundamental challenge for depth estimation and 3D reconstruction due to their violation of Lambertian assumptions, leading to severe geometry degradation in downstream tasks. To address this, we propose a novel geometry-guided preprocessing framework \textbf{GHOST} that leverages visual foundation models to transform transparent regions into opaque, structurally consistent representations without requiring downstream model retraining. Specifically, our pipeline utilizes (1) \textbf{TransDINO} and (2) \textbf{TransDecomp} to disentangle masks and transparency physical properties, while (3) \textbf{DAF-Net} recovers surface normal priors to encode geometric curvature. Subsequently, (4) \textbf{GeoSemTransNet} integrates these multi-modal cues to synthesize a texture-rich opaque RGB image that preserves the transparent object's 3D structure. Extensive experiments demonstrate that our method significantly enhances the accuracy of state-of-the-art depth estimation and reconstruction models on transparent objects by restoring essential photometric cues.