arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoRe-SAM3:面向SAM3裂缝分割的条件语义-视觉调和

CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation

Shipeng Liu, Liang Zhao, Dengfeng Chen

arXiv 2609.05816首次发表:更新:

AI 中文总结

针对SAM3裂缝分割中的漏检与误检问题,提出条件语义-视觉调和(CoRe),以语义为主路径、视觉生成条件残差修正预测,仅增18.914K参数,将平均Crack IoU提升至70.47%。

AI 中文摘要

裂缝分割要求模型在识别目标语义的同时,准确恢复细薄、低对比度且拓扑连续的局部结构。尽管SAM3提供了强大的开放概念分割能力,但其直接应用于裂缝领域仍会遗漏微弱裂缝、激活类似裂缝的背景区域,并产生局部边界误差。我们首先在五个裂缝数据集上诊断了SAM3内部提示条件语义表示与原生视觉表示之间的功能差异。结果表明,语义表示已承载了裂缝预测的大部分任务信息,而视觉表示的效用取决于当前的语义状态。直接结合这两种表示并不能带来一致的性能提升。基于这一发现,我们提出了条件语义-视觉调和(CoRe)。CoRe保留语义预测作为主要决策路径,应用轻量级语义校准来调整目标域决策映射,并利用空间对齐的原生视觉证据生成一个零初始化、有界且正则化的条件残差,以选择性地修正现有预测。在五个域中,CoRe-SAM3将平均Crack IoU从62.34%提升至70.47%,clDice从81.98%提升至89.24%,同时仅引入18.914K个可训练参数。预测转换分析进一步表明,CoRe平均纠正了34.38%的原生错误,而对原生SAM3正确分类像素的损害率仅为0.23%。这些结果表明,基于内部表示功能差异的约束预测修正,为具有强任务特定语义先验的视觉基础模型提供了一种有效且参数高效的目标域适应策略。

英文摘要

Crack segmentation requires a model to recognize target semantics while accurately recovering thin, low-contrast, and topologically continuous local structures. Although SAM3 provides strong open-concept segmentation, its direct application to the crack domain still misses weak cracks, activates crack-like background regions, and produces local boundary errors. We first diagnose the functional differences between the internal prompt-conditioned semantic representation and native visual representation of SAM3 on five crack datasets. The results show that the semantic representation already carries most task information for crack prediction, whereas the utility of the visual representation depends on the current semantic state. Directly combining the two representations does not yield consistent gains. Based on this finding, we propose Conditional Semantic--Visual Reconciliation, termed CoRe. CoRe retains semantic prediction as the primary decision path, applies lightweight semantic calibration to adjust the target-domain decision mapping, and uses spatially aligned native visual evidence to generate a zero-initialized, bounded, and regularized conditional residual that selectively corrects existing predictions. Across five domains, CoRe-SAM3 improves the average Crack IoU from 62.34% to 70.47% and clDice from 81.98% to 89.24%, while introducing only 18.914 K trainable parameters. Prediction-transition analysis further shows that CoRe corrects an average of 34.38% of native errors, with a damage rate of only 0.23% on pixels correctly classified by native SAM3. These results demonstrate that constrained prediction correction based on the functional differences between internal representations provides an effective and parameter-efficient target-domain adaptation strategy for vision foundation models with strong task-specific semantic priors.

CommentsCode will be available at: https://github.com/xauat-liushipeng/CoRe-SAM3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑