arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于杂乱无人机影像中通信塔部件零样本分割的显著性-深度条件设置

Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery

Ali Lesani, Chul Min Yeum, Su-Min Kang

arXiv 2608.25435首次发表:更新:

发表机构

University of Waterloo; Soongsil University(滑铁卢大学; 崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无人机影像通信塔部件零样本分割的背景干扰问题,提出显著性-深度前景条件策略,集成至Grounded-SAM与SAM 3,在TOW-300数据集上提升了实例分割性能并减少误报。

AI 中文摘要

对无人机(UAV)影像中通信塔部件进行细粒度分割是自动化巡检的关键,但因实例级标注有限,难以开发任务特定模型。零样本分割模型是有前景的替代方案,但在杂乱场景中,视觉相似的背景结构会干扰部件定位,导致漏检和误报。我们提出一种与模型无关的显著性-深度前景条件策略,结合基于外观的显著性与单目相对深度,构建粗略的塔先验以抑制无关内容。我们将该模块与Grounded-SAM、SAM 3集成,得到SD-Grounded-SAM和SD-SAM 3。SD-Grounded-SAM在生成掩码前进一步应用几何与深度感知的框优化,而SD-SAM 3依赖SAM 3的内部设置。在含340张通信塔无人机影像的TOW-300数据集上,我们的策略提升了两个基线:SD-SAM 3实现了最强的实例分割性能,SD-Grounded-SAM则产生更少的误报。 ablation实验证实了显著性、深度和框优化的互补增益,提升了杂乱场景中的鲁棒性。

英文摘要

Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually similar background structures interfere with component localization, causing missed instances and false positives. We propose a model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content. We integrate this module with Grounded-SAM and SAM 3, yielding SD-Grounded-SAM and SD-SAM 3. SD-Grounded-SAM further applies geometric and depth-aware box refinement before mask generation, while SD-SAM 3 relies on SAM 3's internal setup. On TOW-300, a dataset of 340 communication-tower UAV images, our strategy improves both baselines: SD-SAM 3 achieves the strongest instance-segmentation performance, while SD-Grounded-SAM produces fewer false positives. Ablations confirm complementary gains from saliency, depth, and box refinement, improving robustness in cluttered scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑