发表机构
Department of Information Science and Applied Artificial Intelligence, Bar-Ilan University, Ramat Gan, Israel(信息科学与应用人工智能系,巴伊兰大学,拉马特甘,以色列)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究谷歌MedGemma-4B在安全护栏方面的问题,构建全因子基准测试,通过多种攻击方式和法官编码量化其表现,发现重新解释请求的框架能提高攻击成功率,且稳健性受主题主导,呼吁为开放医学模型加强部署时护栏。
AI 中文摘要
开放权重的医学语言模型越来越多地被用作面向患者和临床医生支持应用程序的基础。其模型卡禁止特定行为,如推荐确切药物剂量、做出明确诊断、开处方、判定药物相互作用以及建议可跳过紧急护理等,但模型卡描述的是预期行为而非稳健行为。我们对MedGemma-4B的这种差距进行了量化,它在无需复杂技术的攻击下表现不佳。我们构建了一个全因子基准,包括5个受保护行为概念、50个确定性模板问题、6种易于访问的攻击方式和3次重复(共4500个生成),通过Ollama在默认采样下本地运行模型,并由三名独立法官(一个语言模型法官、一个透明正则表达式法官和一个自然语言推理法官)对每个回答(拒绝/回避/遵守)进行编码。在主要的语言模型法官下,总体攻击成功率(ASR,编码为遵守的比例)为38.0%。将请求重新解释为合法的两种框架占主导地位:将问题重新表述为“医学委员会考试”项目使ASR从29.0%的基线提高到53.1%(提高24.0个百分点),诉诸所谓医生的权威将其提高到43.7%(提高14.7个百分点);粗略的指令覆盖前缀没有显著影响。稳健性受主题主导:药物相互作用护栏几乎不存在(ASR为83.2%),而紧急推迟护栏很强(4.7%),诉诸权威框架是唯一突破它的攻击。我们报告了威尔逊置信区间、聚类自举效应大小、聚类稳健逻辑回归、 Cochr an's Q、每种方式的 McNemar 检验以及法官间的可靠性(Fleiss' kappa = 0.26);绝对ASR取决于法官,而攻击和主题的排序则不然。我们的发现促使为开放医学模型制定更强的部署时护栏。
英文摘要
Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recommending exact drug dosages, issuing definitive diagnoses, prescribing treatments, adjudicating drug-drug interactions, and advising that emergency care can be skipped -- yet a model card describes intended behavior, not robust behavior. We quantify that gap for MedGemma-4B-it under attacks that require no technical sophistication. We build a fully factorial benchmark of 5 guarded-behavior concepts x 50 deterministically templated questions x 6 lay-accessible attack manners x 3 repetitions (4,500 generations), serve the model locally through Ollama under default sampling, and code every response refuse/hedge/comply with three independent judges (an LLM judge, a transparent regex judge, and an NLI-entailment judge). Under the primary LLM judge the overall Attack Success Rate (ASR, the fraction coded comply) is 38.0%. The two framings that reinterpret the request as legitimate dominate: recasting a question as a "medical board exam" item raises ASR from a 29.0% baseline to 53.1% (+24.0 points), and an appeal to an alleged doctor's authority raises it to 43.7% (+14.7); crude instruction-override prefixes have no significant effect. Robustness is dominated by topic: the drug-interaction guardrail is nearly absent (83.2% ASR) while the emergency-deferral guardrail is strong (4.7%) -- and the authority framing is the only attack that breaches it. We report Wilson confidence intervals, cluster-bootstrap effect sizes, a cluster-robust logistic regression, Cochran's Q, per-manner McNemar tests, and inter-judge reliability (Fleiss' kappa = 0.26); absolute ASR is judge-dependent while the ordering of attacks and topics is not. Our findings motivate stronger deployment-time guardrails for open medical models.