发表机构
Universiti Malaya; BRAC University(马来亚大学; BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MedPRESS是衡量大语言模型患者压力诱导型医疗谄媚行为的多轮基准,含600个五轮医学对话,评估20类LLM发现其易在压力下不安全附和,反谄媚提示仅部分改善,凸显医疗LLM需抗对话压力的缺口。
AI 中文摘要
大语言模型(LLM)越来越多地被用于提供健康相关建议。现有研究通过静态问题而非面向患者的压力对话来衡量其安全性。我们推出MedPRESS,这是一个用于衡量LLM中患者压力诱导型谄媚行为的多轮基准测试。MedPRESS包含600个基于医学依据的五轮对话,涵盖三个场景类别:药物与治疗需求、个人健康自我护理,以及症状分诊与护理抵抗。每个对话以健康查询开始,通过个人经历、社会证明、外部证据主张和直接对抗挑战逐步升级。我们使用结构化评判和以安全为重点的指标,评估了20个属于通用、医疗领域、轻量、大型、开放权重及专有类别的LLM。结果显示,在反复的患者压力下,模型频繁转向不安全的一致回应,且不同模型类别、模型规模和提示类型之间存在显著差异。反谄媚提示提升了部分模型的鲁棒性,但并未消除不安全的一致回应。MedPRESS凸显了医疗LLM评估中的一个关键缺口:仅具备安全的医学知识是不够的,除非模型能在对话压力下保持该知识。
英文摘要
Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than pressured patient-facing conversations. We introduce MedPRESS, a multi-turn benchmark for measuring patient-pressure-induced sycophancy in LLMs. MedPRESS contains 600 medically grounded five-turn dialogues across three scenario families: medication and treatment demand, personal health self-care, and symptom triage and care resistance. Each dialogue begins with a health query and escalates through personal experience, social proof, external evidence claims, and direct adversarial challenge. We evaluate 20 LLMs across general, medical-domain, lightweight, large, open-weight, and proprietary families using structured judging and safety-focused metrics. Results show that models frequently shift toward unsafe agreement under repeated patient pressure, with substantial variation across model families, model scale, and prompt type. Anti-sycophancy prompting improves robustness for several models, but does not eliminate unsafe agreement. MedPRESS highlights a critical gap in medical LLM evaluation: safe medical knowledge is not enough unless models can maintain it under conversational pressure.
Comments27 pages, 10 figures. Both authors contributed equally