arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.06846cs.CLcs.AI

MedDialBench:在参数对抗患者行为下评估LLM诊断鲁棒性的基准测试

MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors

  • Shanda Group(盛大集团)

机构由 AI 辅助整理,请以论文原文为准。

Xiaotian Luo, Xun Jiang, Jiangcheng Wu

更新

AI总结:

MedDialBench通过控制患者行为维度,评估LLM在不同行为下的诊断鲁棒性,发现信息污染对准确率的影响大于信息缺失,且模型存在不同的脆弱性特征。

AI中文摘要:

交互式医疗对话基准测试显示,当与不合作患者交互时,LLM的诊断准确性显著下降,但现有方法要么缺乏严重程度分级或具体案例支撑,要么将患者不合作简化为单一轴。我们引入MedDialBench,该基准测试能够控制并分析不同患者行为维度对LLM诊断鲁棒性的影响。它将患者行为分解为五个维度——逻辑一致性、健康认知、表达风格、披露和态度,每个维度都有分级严重程度和案例特定的行为脚本。这种受控的因子设计允许进行分级敏感性分析、剂量反应分析和跨维度交互检测。评估五种前沿LLM在7225次对话(85个案例x17种配置x5种模型)中,发现信息污染(制造症状)导致的准确率下降比信息缺失(隐瞒信息)大1.7-3.4倍,且制造症状是唯一在所有五种模型中均具有统计显著性的配置(McNemar p < 0.05)。在六个维度组合中,制造是唯一导致超加性交互的因素:所有涉及制造的配对产生O/E比为0.70-0.81(35-44%的合格案例在组合下失败,尽管在每个维度单独时成功),而所有非制造配对显示纯加性效应(O/E ~ 1.0)。询问策略调节缺失但不调节污染:彻底提问恢复被隐瞒的信息,但无法补偿伪造输入。模型表现出不同的脆弱性特征,最坏情况下降范围从38.8到54.1个百分点。

英文摘要:

Interactive medical dialogue benchmarks have shown that LLM diagnostic accuracy degrades significantly when interacting with non-cooperative patients, yet existing approaches either apply adversarial behaviors without graded severity or case-specific grounding, or reduce patient non-cooperation to a single ungraded axis, and none analyze cross-dimension interactions. We introduce MedDialBench, a benchmark enabling controlled, dose-response characterization of how individual patient behavior dimensions affect LLM diagnostic robustness. It decomposes patient behavior into five dimensions -- Logic Consistency, Health Cognition, Expression Style, Disclosure, and Attitude -- each with graded severity levels and case-specific behavioral scripts. This controlled factorial design enables graded sensitivity analysis, dose-response profiling, and cross-dimension interaction detection. Evaluating five frontier LLMs across 7,225 dialogues (85 cases x 17 configurations x 5 models), we find a fundamental asymmetry: information pollution (fabricating symptoms) produces 1.7-3.4x larger accuracy drops than information deficit (withholding information), and fabricating is the only configuration achieving statistical significance across all five models (McNemar p < 0.05). Among six dimension combinations, fabricating is the sole driver of super-additive interaction: all three fabricating-involving pairs produce O/E ratios of 0.70-0.81 (35-44% of eligible cases fail under the combination despite succeeding under each dimension alone), while all non-fabricating pairs show purely additive effects (O/E ~ 1.0). Inquiry strategy moderates deficit but not pollution: exhaustive questioning recovers withheld information, but cannot compensate for fabricated inputs. Models exhibit distinct vulnerability profiles, with worst-case drops ranging from 38.8 to 54.1 percentage points.

补充信息

↑