arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究多智能体医疗系统在临床推理中对人为干预的脆弱性

Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning

Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang, Abhiram Gudimella, Kevin Zhu

arXiv 2609.02191首次发表:更新:

发表机构

Stanford University; Jordan High School; UC San Diego; Winchester High School; James Logan High School; Rutgers University; Foothill–De Anza College; Fordham University; Chapman University; UC Berkeley(斯坦福大学; 乔丹高中; 加利福尼亚大学圣迭戈分校; 温彻斯特高中; 詹姆斯·洛根高中; 罗格斯大学; 福希尔-德安扎学院; 福特汉姆大学; 查普曼大学; 加利福尼亚大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多智能体医疗系统,借助MedQA数据集分析人为干预对其临床推理的影响,发现正确干预可提升诊断准确率、错误干预会降低性能,还揭示了模拟智能体与现实临床的认知偏见相似性,为提升系统诊断鲁棒性提供思路。

AI 中文摘要

人为在故障点进行干预可改变多智能体医疗系统的诊断准确率。我们将故障点定义为AI智能体对话中,智能体推理最易受外部影响的时刻。本研究使用MedQA数据集,分析模拟的医患对话,以衡量干预如何改变推理过程和准确率。正确的干预方法可使基线诊断准确率提升高达40%,而错误或与偏见相关的干预则会使性能下降高达6%,并增加诊断偏差和不确定性。除性能变化外,我们的分析还发现模拟智能体环境中的认知偏见与现实临床实践存在行为相似性,例如过早闭合和易受误导线索影响。总体而言,这些发现表明,识别人为干预的故障点并加以引导,可能为提升多智能体医疗系统的诊断鲁棒性提供一种机制。

英文摘要

Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this study analyzed simulated doctor-patient conversations to measure how interventions shifted reasoning and accuracy. Correct intervention methods showed an improvement in baseline diagnostic accuracy of up to 40%, while incorrect or bias-related interventions degraded performance by up to 6% and increased diagnostic drift and uncertainty. Beyond performance changes, our analysis revealed behavioral similarities between cognitive biases in simulated agent environments and real-world clinical practice. Examples included premature closure and susceptibility to misleading cues. Overall, these findings demonstrate that identifying and guiding fault points with human interventions may provide a mechanism for improving diagnostic robustness in multi-agent medical systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑