发表机构
Xi’an University of Architecture and Technology(西安建筑科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对裂缝分割中现有方法需特定注释训练及基础模型掩码不适用于薄裂缝问题,提出语义边缘响应解码(SERD),无需注释微调,实验表明其性能优于原生SAM3及其他对比方法,证明内部连续响应更具可转移性。
AI 中文摘要
裂缝分割对基础设施检查和结构健康评估至关重要,但现有高性能方法通常需要特定任务的像素级注释和训练。文本提示视觉基础模型可实现零样本部署,但其最终掩码提议不适用于薄、碎片化和低对比度裂缝。研究发现SAM3解码器内语言条件语义响应保留的裂缝证据更连续完整,据此提出语义边缘响应解码(SERD),无需注释或微调。在六个公共数据集上的实验表明,SERD持续优于原生SAM3,平均裂缝IoU达61.14%,比SAM3高4.63个百分点。结果表明,对于薄且不紧凑目标,内部连续响应比基础模型的最终掩码提供更可转移的接口。
英文摘要
Reliable crack measurement is essential for infrastructure condition assessment, yet existing image-based approaches typically depend on pixel-wise annotations, task-specific segmentation training, and mask-based geometric measurement, making cross-scene deployment costly and sensitive to segmentation errors. We identify an output-interface mismatch in SAM3: its prompt-conditioned semantic response preserves crack evidence that is often suppressed or spatially distorted in the final candidate masks. Across six public crack datasets, the internal response achieves 82.66% average crack-pixel recall, compared with 74.66% for the retained SAM3 proposals, with an average mismatch ratio of 8.73%. Based on this observation, we propose Semantic-Edge Response Decoding (SERD) to calibrate the semantic response using a fixed Sobel structural field, and further develop SERD-DQ, a training-free framework that directly estimates crack centerline and transverse geometry from the continuous decoded response without generating an intermediate predicted mask. Experiments verify both segmentation fidelity and direct geometric measurement against manually established pixel-level references. Compared with native SAM3 mask-based quantification, SERD-DQ reduces width MAE from 6.072 to 5.547 pixels, length relative error from 22.245% to 17.355%, and area relative error from 33.921% to 26.228%, while achieving a latent geometry recovery rate of 0.394. The results indicate that continuous semantic-edge responses provide a more reliable interface for training-free crack quantification than conventional mask-mediated measurement.
CommentsSubmitted to Elsevier