发表机构
MAA Consultants Co., Ltd.; University of Phayao; Thailand Institute of Scientific and Technological Research (TISTR)(MAA咨询有限公司; 帕尧大学; 泰国科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将LoRA应用于SAM3,提出无需额外标注的监督流程缓解失效模式,在两个结构缺陷数据集上大幅提升了多类分割性能,实现参数高效适配。
AI 中文摘要
SAM3等可提示分割基础模型接受开放词汇文本概念,返回所有匹配的实例,但对最需要该技术的机构而言,对其进行全微调以适配专业领域的计算成本过高。本研究将低秩适配(LoRA)应用于SAM3,用于多类结构缺陷分割,探究该模型能否通过常规标注进行监督,以及由此获得的效率提升是否可跨数据集迁移。两项方法学贡献如下:其一,提出一种监督流程,该流程直接从COCO风格的类别标注实例分割数据训练概念可提示模型,将类别名称本身作为提示,无需提示模板、同义词扩展或学习到的类别嵌入。其二,识别并缓解了该场景特有的失效模式:由于常规标注文件仅产生正提示,模型的存在预测与文本条件解耦,退化为对任何提示都做出响应,这种崩溃在仅基于正提示计算的所有指标中都不可见; exhaustive hard-negative prompting(对图像中不存在的每个数据集类别作为零检测查询)可解决该问题,且无需额外标注成本。在相同协议下比较了两种适配器位置,分别更新模型0.121%和1.341%的参数;在自建隧道衬砌数据集上,像素交并比从0.017提升至0.338,实例级召回率从0.375提升至0.672;在独立的公开结构缺陷数据集(Structural Defects Dataset)上,分别从0.017提升至0.855、从0.574提升至1.000。在两个数据集的10项指标上,改进方向一致,且每类别的最大提升恰好出现在零样本能力缺失的地方。
英文摘要
Promptable segmentation foundation models such as SAM3 accept an open-vocabulary text concept and return every instance matching it, but adapting them to a specialized domain by full fine-tuning is computationally prohibitive for the organizations that would benefit most. This study applies Low-Rank Adaptation (LoRA) to SAM3 for multi-class structural defect segmentation and examines both how such a model can be supervised from conventional annotation and whether the resulting efficiency gain transfers across datasets. Two contributions are methodological. First, we describe a supervision procedure that trains a concept-promptable model directly from COCO-style class-labeled instance segmentation by using the category name itself as the prompt, requiring no prompt templates, no synonym expansion, and no learned class embeddings. Second, we identify and mitigate a failure mode specific to this setting: because a conventional annotation file yields positive prompts exclusively, the model's presence prediction decouples from the text condition and degenerates into responding to any prompt, a collapse that is invisible to every metric computed on positive prompts alone. Exhaustive hard-negative prompting, in which every dataset category absent from an image is issued as a zero-detection query, addresses this at no annotation cost. Two adapter placements were compared under an identical protocol, updating 0.121% and 1.341% of model parameters. On a purpose-built tunnel lining dataset, pixel intersection-over-union improved from 0.017 to 0.338 and instance-level recall from 0.375 to 0.672; on the independent public Structural Defects Dataset, from 0.017 to 0.855 and from 0.574 to 1.000. Improvements were directionally consistent across ten metrics on both datasets, and the largest per-category gains occurred precisely where zero-shot competence was absent.