arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DART:面向外科视觉基础模型的以深度为目标的预训练

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri

arXiv 2609.04555首次发表:更新:

发表机构

Vanderbilt University; Intuitive Surgical, Inc.(范德堡大学; 直觉外科公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DART是基于DINOv2的RGB-D预训练方案,利用伪标签深度作为预训练目标,在8项外科基准测试中优于各类基线,可强化外科视觉基础模型的表示能力且不增加推理成本。

AI 中文摘要

视觉基础模型(VFMs)在外科手术等数据稀缺领域具有重要价值,单个预训练骨干网络可为诸多下游任务提供丰富表示。然而,主流自监督预训练范式仅使用RGB图像,未利用 readily available 的互补信号(如深度图),这对外科手术领域而言是明显的机会损失——自然图像VFMs的迁移效果较差,而该领域的场景几何信息丰富且具有高价值。当前已有成熟的现成模型可为任意图像语料生成伪标签化的密集深度,因此我们假设可将此类信号融入预训练过程以学习更优的表示。我们提出了DART,一种基于DINOv2的RGB-D预训练方案,仅做了简单修改:对带掩码的iBOT补丁应用像素级深度重建目标,由伪标签深度进行监督。深度仅在预训练阶段使用,微调与推理阶段仍仅使用RGB。我们发现该像素级重建头可提升表示质量,而非破坏其性能。进一步研究表明,编码场景几何信息的深度作为目标比Canny边缘等其他密集信号更有效,证实性能提升源于深度本身,而非仅增加的监督信号。在涵盖分割、深度估计和图像级识别的8个外科基准测试中,DART的表现优于自然图像基线和领域内基线,包括在相同数据上训练的普通DINOv2,既提升了密集预测性能,也增强了图像级理解能力。更广泛而言,DART表明,免费可得的几何伪标签可在无需额外标签或增加推理成本的情况下强化基础模型预训练,为构建更强的外科领域骨干网络指明了方向。

英文摘要

Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretraining paradigm uses only RGB images, leaving readily available complementary signals, such as depth maps, unused. This is a particular missed opportunity in surgery, where natural-image VFMs transfer poorly while the scene geometry is rich and informative. With strong off-the-shelf models now able to produce pseudo-labeled dense depth for any image corpus, we hypothesize that such signals can be folded into pretraining to learn better representations. We present DART, an RGB-D pretraining recipe that builds on DINOv2 with a simple modification: a pixel-space depth reconstruction objective applied to masked iBOT patches, supervised by pseudo-labeled depth. Depth is used only during pretraining, so fine-tuning and inference remain RGB-only. We find that this pixel-level reconstruction head improves representation quality rather than disrupting it. We further show that depth, which encodes scene geometry, is more effective as a target than alternative dense signals such as Canny edges, confirming that the gains stem from depth rather than added supervision alone. Across eight surgical benchmarks spanning segmentation, depth estimation, and image-level recognition, DART outperforms both natural-image and in-domain baselines, including a vanilla DINOv2 trained on identical data, improving dense prediction while also strengthening image-level understanding. More broadly, DART shows that freely available geometric pseudo-labels can strengthen foundation model pretraining without extra labels or added inference cost, pointing toward stronger backbones for surgery.

CommentsAccepted to BMVC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑