基于强化学习的医疗问卷系统的对抗性攻击:输入级扰动策略与医学约束验证
Adversarial Attacks on Reinforcement Learning-based Medical Questionnaire Systems: Input-level Perturbation Strategies and Medical Constraint Validation
AI总结:
本研究针对基于强化学习的医疗问卷系统,提出输入级对抗扰动策略,并通过247条医学约束验证,发现其攻击成功率最高达64.70%,揭示系统存在显著安全漏洞。
AI中文摘要:
基于强化学习(RL)的医疗问卷系统在医疗场景中展现出巨大潜力。然而,其安全性和鲁棒性仍未得到解决。本研究对对抗性攻击方法进行了全面评估,以识别和分析其潜在漏洞。我们将诊断过程建模为马尔可夫决策过程(MDP),其中状态是患者的回答和未提出的问题,动作是提问或做出诊断。我们实现了六种主流攻击方法,包括快速梯度符号法(FGSM)、投影梯度下降(PGD)、Carlini & Wagner攻击(C&W)、基本迭代法(BIM)、DeepFool和AutoAttack,每种方法使用七个epsilon值。为确保生成的对抗样本在临床上具有合理性,我们开发了一个包含247条医学约束的综合医学验证框架,涵盖生理界限、症状相关性和条件性医学约束。我们在生成临床合理的对抗样本方面达到了97.6%的成功率。我们在国家健康访谈调查(NHIS)数据集(https://www.cdc.gov/nchs/nhis/)上进行了实验,该数据集包含182,630个样本,用于预测参与者4年死亡率。我们在arXiv:2004.00994中提出的AdaptiveFS框架上评估了我们的攻击。结果表明,对抗性攻击能显著影响诊断准确性,攻击成功率从33.08%(FGSM)到64.70%(AutoAttack)不等。我们的工作表明,即使在对输入施加严格医学约束的情况下,此类基于RL的医疗问卷系统仍表现出显著漏洞。
英文摘要:
RL-based medical questionnaire systems have shown great potential in medical scenarios. However, their safety and robustness remain unresolved. This study performs a comprehensive evaluation on adversarial attack methods to identify and analyze their potential vulnerabilities. We formulate the diagnosis process as a Markov Decision Process (MDP), where the state is the patient responses and unasked questions, and the action is either to ask a question or to make a diagnosis. We implemented six prevailing major attack methods, including the Fast Gradient Signed Method (FGSM), Projected Gradient Descent (PGD), Carlini & Wagner Attack (C&W) attack, Basic Iterative Method (BIM), DeepFool, and AutoAttack, with seven epsilon values each. To ensure the generated adversarial examples remain clinically plausible, we developed a comprehensive medical validation framework consisting of 247 medical constraints, including physiological bounds, symptom correlations, and conditional medical constraints. We achieved a 97.6% success rate in generating clinically plausible adversarial samples. We performed our experiment on the National Health Interview Survey (NHIS) dataset (https://www.cdc.gov/nchs/nhis/), which consists of 182,630 samples, to predict the participant's 4-year mortality rate. We evaluated our attacks on the AdaptiveFS framework proposed in arXiv:2004.00994. Our results show that adversarial attacks could significantly impact the diagnostic accuracy, with attack success rates ranging from 33.08% (FGSM) to 64.70% (AutoAttack). Our work has demonstrated that even under strict medical constraints on the input, such RL-based medical questionnaire systems still show significant vulnerabilities.