arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型护栏的幻象:人工智能辅助医疗记录操纵的案例研究

The Mirage of LLM Guardrails: A Case Study in AI-Assisted Medical Note Manipulation

Davis Yadav, Amulya Yadav

arXiv 2607.24859首次发表:更新:

AI 中文总结

研究大语言模型在医疗记录操纵场景下内置护栏的可靠性,通过开发操纵流程、实证评估、利用多种指标评估及用户研究,揭示其弱点,为大语言模型在医疗领域的护栏设计、安全政策及伦理等方面提供参考。

AI 中文摘要

大语言模型(LLMs)在医疗环境中的快速部署,使得其针对恶意查询的内置护栏的可靠性成为一个具有紧迫实际影响的问题。然而,这些机制在医疗背景下抵御蓄意滥用的稳健性仍知之甚少。本文以人工智能辅助医疗记录操纵为具体案例进行实证研究。有四项新贡献:开发了可重复的操纵流程,用公开可用的医疗记录模板,通过商业LLMs跨模型家族、输入格式和提示措辞替换患者姓名等生成定制操纵记录;对LLM护栏在医疗记录操纵方面的稳健性进行系统实证评估,发现当代商业LLM护栏存在重大弱点和不一致性;利用自动化指标和基于人工标注的指标评估操纵请求的正确性;进行用户研究评估操纵医疗记录的可信度,发现最佳操纵记录对人工评分者来说在视觉上与原始文档难以区分。最后讨论了对LLMs中负责任的护栏设计、人工智能安全政策以及在医疗环境中部署LLMs的更广泛伦理的影响。

英文摘要

The rapid deployment of large language models (LLMs) in healthcare settings makes the reliability of their built-in guardrails against malicious queries a question of urgent practical consequence. Yet the robustness of these mechanisms against deliberate misuse (in the healthcare context) remains poorly understood. In this paper, we investigate this question empirically, using AI-assisted medical note manipulation as a concrete case study. We make four novel contributions. First, we develop a reproducible manipulation pipeline that takes publicly available seed medical note templates and use commercial LLMs to produce customized manipulated notes by substituting patient names, provider identities, dates, and medical conditions across multiple model families, input formats, and prompt phrasings. Second, we conduct a systematic empirical evaluation of LLM guardrail robustness for medical note manipulation. Our experimental results reveal substantial weaknesses and inconsistencies in contemporary commercial LLM guardrails, including low refusal rates for several model families. Third, we utilize a combination of automated metrics and human annotation-based metrics to assess the correctness of requested manipulations. Fourth, we conduct a user-study to assess the believability of manipulated medical notes, finding that the best manipulations are visually indistinguishable from original documents to human raters. Finally, we discuss implications for responsible guardrail design in LLMs, AI safety policies, and the broader ethics of deploying LLMs in healthcare settings.

Comments10 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑