正确诊断,装饰性推理:医学思维链的扰动审计
Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
- Eindhoven University of Technology(埃因霍温理工大学)
- Dana-Farber Cancer Institute(达纳-法伯癌症研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究通过医学扰动审计方法,对14个大语言模型在四个医学问答基准上进行测试,发现医学思维链多为装饰性内容,为医学CoT的忠实性审计提供了可复用标准。
中文摘要 AI 辅助
临床医生将思维链(CoT)推理过程视为医疗推理的证据,但这一可见的思维链是否真的起到该作用却很少被验证。通用领域的CoT忠实性探测忽略了医疗成本,而医疗大语言模型(LLM)评估则将思维链视为黑箱。本研究通过医学扰动审计填补这一空白:采用包含30种算子的工具集,以具有临床动机的算子(严重程度反转、否定翻转、人口统计学特征交换、证据消融)编辑思维链和问题,结合思维链更新与答案翻转的联合分析,按失败模式对每个模型进行分类。在四个医学问答基准上对14个LLM应用该方法后,三项独立测试得出一致结果:在具有临床意义的破坏性编辑下,全面板的思维链解耦率(CDR;思维链未记录编辑且答案未翻转)为72.9%,思维链损坏不会改变准确率,移除CoT提示不会降低准确率。两名持证临床医生重新标注了N=197个扰动问题,其中98.5%的标准答案仍合理。该模式在医疗和推理微调及不同规模模型中均成立;在无法获取思维链文本的闭源模型层级,答案侧信号与相同的解耦情况一致。本研究的框架和CDR为审计医学CoT是否忠实或仅为装饰性内容提供了可复用的标准。
英文摘要
Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely tested. General-domain CoT-faithfulness probes ignore clinical cost, and medical LLM evaluations treat the chain as a black box. We close this gap with a medical perturbation audit: a 30-operator battery edits both the chain and the question with clinically motivated operators (severity reversal, negation flip, demographic swap, evidence ablation), paired with a chain-update times answer-flip joint analysis that classifies each model by its failure mode. Applied to 14 LLMs on four medical QA benchmarks, three independent tests converge: the Chain-Decoupling Rate (CDR; chain does not register the edit and the answer does not flip) is 72.9% panel-wide on clinically meaningful destructive edits, chain corruption leaves accuracy unchanged, and removing CoT prompting does not reduce accuracy. Two board-certified clinicians re-annotate N=197 perturbed questions; 98.5% leave the gold defensible. The pattern holds across medical and reasoning fine-tuning and scale; on the closed-source tier, where the chain text is unavailable, the answer-side signals are consistent with the same decoupling. Our framework and CDR provide a reusable yardstick for auditing whether medical CoT is faithful or merely documentation.