arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14290cs.CRcs.LG

融合频谱签名与激活聚类用于医疗影像模型后门检测:方法、实现与评估

Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models: Method, Implementation, and Evaluation

发表机构坎伯兰大学
查看机构详情
  • University of the Cumberlands(坎伯兰大学)

机构由 AI 辅助整理,请以论文原文为准。

Suresh Tamang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出融合频谱签名与激活聚类的后门检测流水线,在医学影像基准上达到AUROC≥0.99,但在CIFAR-10高投毒率下融合失效,并给出开源实现与NIST/MITRE术语报告。

中文摘要 AI 辅助

机器学习模型日益部署于医疗影像流程中用于诊断支持,针对这些模型的训练时攻击是一个被明确指出的行业级关切:医疗行业指南将模型投毒和对抗性攻击列为需要专门防御的威胁,而联邦政策则指示将扩展的AI漏洞检测工具提供给关键基础设施运营商(如乡村医院)。频谱签名分析和激活聚类是两种成熟的后门检测方法,通常作为独立基线进行评估,但它们的输出通常不进行组合,且在医学影像基准上的检测性能报告相对于自然图像设置仍然稀少。本文贡献三方面内容:一种分数级融合规则,将逐类频谱排序与激活聚类标志组合为单一的逐样本投毒分数及模型级一致性统计量;所得到的八阶段流程的开源实现;以及针对公共医学影像基准和CIFAR-10的合成投毒变体,在四个投毒率(0%、1%、5%、10%)下各五个随机种子进行的评估,比较各检测器单独使用与融合后的效果。在医学基准上,融合检测器在每个非零投毒率下均达到AUROC≥0.99。在CIFAR-10上,融合并非总是有益:在10%投毒率下,激活聚类的真阳性率降至0.000,频谱AUROC独立退化至接近随机(0.545),尽管97.2%的攻击成功率确认后门已完全植入。融合分数是两种信号的加权组合,继承了这一联合失败。检测输出以NIST AI RMF Measure功能和MITRE ATLAS术语表达,因此结果以安全和合规团队已使用的词汇报告。

英文摘要

Machine learning models are increasingly deployed in healthcare imaging pipelines for diagnostic support, and training-time attacks against them are a named sector-level concern: healthcare-sector guidance identifies model poisoning and adversarial attacks as threats requiring dedicated defenses, while federal policy directs expanded AI vulnerability-detection tooling to critical infrastructure operators such as rural hospitals. Spectral signature analysis and activation clustering are two established backdoor detection methods routinely evaluated as independent baselines, but their outputs are not ordinarily combined, and reported detection performance on medical imaging benchmarks remains sparse relative to the natural-image setting. This paper contributes three things: a score-level fusion rule combining per-class spectral ranking with activation-clustering flags into a single per-sample poisoning score and a model-level agreement statistic; an open-source implementation of the resulting eight-stage pipeline; and an evaluation of that pipeline against synthetically poisoned variants of a public medical imaging benchmark and CIFAR-10 at four poisoning rates (0%, 1%, 5%, 10%) over five seeds each, measuring each detector alone against the fusion. On the medical benchmark, the fused detector reaches AUROC >= 0.99 at every nonzero poisoning rate tested. On CIFAR-10, fusion does not uniformly help: at 10% poisoning, activation clustering's true-positive rate collapses to 0.000 and spectral AUROC independently degrades to near-chance (0.545), despite a 97.2% attack success rate confirming the backdoor was fully installed. The fused score, a weighted combination of both signals, inherits this joint failure. Detection output is expressed in NIST AI RMF Measure-function and MITRE ATLAS terms, so findings are reported in the vocabulary security and compliance teams already use.

补充信息

↑