arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38811cs.CVcs.LG

DCM-SAM:面向NPU部署的增材制造缺陷分割的缺陷条件LoRA专家混合模型

DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation

  • West Virginia University(西弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Md Mushfiqur Rahaman, Md Mahedi Hasan, Imtiaz Ahmed, Srinjoy Das

AI总结:

提出DCM-SAM,一种缺陷条件LoRA专家混合模型,在合成数据上训练,仅更新4.4%参数,在XCT-SAM基准和真实NIST扫描上超越基线,并通过注意力重写实现NPU上FP16高效部署。

AI中文摘要:

金属增材制造零件通过X射线计算机断层扫描进行检测,其中标注数据稀缺,关键的孔隙和夹杂物仅跨越几个像素,且检测必须在机器上进行。我们提出DCM-SAM,一种缺陷条件的自适应LoRA专家混合模型:一个冻结的Segment Anything骨干网络为每个缺陷类别携带独立的Conv-LoRA专家库和掩码解码器,每个类别在各自的训练过程中,仅使用合成切片,无需提示,仅更新4.4%的参数。在XCT-SAM报告的基准上,DCM-SAM在ViT-B骨干网络下对两个类别均优于所有基线(对比其ViT-H),并在未见任何真实图像的情况下,在真实NIST扫描上达到64.2%的孔隙IoU。部署随后暴露了适应工作很少衡量的内容:在Qualcomm Hexagon NPU上,ViT-H和ViT-L可编译但无法在1024x1024图像分辨率下分配内存,因为激活而非权重超出设备上限,且量化权重无济于事。仅ViT-B可运行,但适应后的编码器在原始编码器成功之处分配失败,直到对注意力进行数值上等效的重写,使完整的DCM-SAM以FP16在1024x1024下运行,无算子回退至CPU,掩码与FP32参考的像素差异在0.01%以内。代码:此https URL。

英文摘要:

Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, each trained in its own pass, without prompts, on synthetic slices alone, updating only 4.4% of the parameters. On benchmarks that XCT-SAM reports, DCM-SAM improves on every baseline for both classes from a ViT-B backbone against their ViT-H, and reaches 64.2% pore IoU on real NIST scans having seen no real images during training. Deployment then exposes what adaptation work rarely measures: on a Qualcomm Hexagon NPU, ViT-H and ViT-L compile yet cannot allocate at 1024x1024 image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help. ViT-B alone runs, but the adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at 1024x1024, with no operator falling back to the CPU, masks within 0.01% of pixels of the FP32 reference. Code: https://github.com/MushfiqShovon/DCM-SAM.

补充信息

↑