发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对医学问答中误导性上下文的机制,在MedMisBench基准上测试三类模型,发现断言型误导线索更易干扰模型,仅开放推理轨迹可可靠监测错误决策。
AI 中文摘要
大型语言模型现在能够以专家级性能回答医学问题。然而,这些系统所依据的上下文可能具有误导性,而误导性上下文会破坏模型的医学判断。为了理解误导性上下文如何破坏这种判断,我们研究了模型对上下文的敏感性、对上下文的披露、推理被破坏的机制以及决策的可监测性。在 MedMisBench(一个由临床医生审核的包含8627个问题的问答基准)的医学推理子集上,我们注入了两种类型的误导性上下文线索:伪造证据和单纯断言。我们测试了三个推理模型,其中两个会暴露其完整推理轨迹,另一个前沿模型仅暴露其响应。所有三个模型对断言的敏感性都高于对伪造证据的敏感性,采纳断言答案的频率高出10至27个百分点。误导性线索在81%至98%的轨迹中被披露,但仅在7%至90%的响应中被披露,且断言的披露频率低于基于证据的线索。从未披露的推理轨迹中重新采样显示,两种线索破坏推理的方式不同:证据早期进入并累积,而断言在接近末尾时重定向结论。当LLM监测器在阅读开放模型的轨迹并获得指导时,能以5%的假阳性率捕获78%的错误决策,而从任何响应中捕获的比例最多为32%。模型最易受影响的误导性上下文披露最少,且仅能从前沿提供商 withheld 的开放推理轨迹中被可靠捕获。
英文摘要
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.
Comments25 pages, 10 figures. Submitted to ML4H 2026