发表机构
Guangzhou University; Institute of Automation, Chinese Academy of Sciences(广州大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对医疗智能体评估忽略中间推理质量的问题,提出MedTraj框架,通过多维度评分、受控错误注入等优化,在多个医学数据集上提升了推理连贯性、正确性并降低了幻觉比例。
AI 中文摘要
医疗人工智能智能体的评估目前主要以答案为中心,仅评估最终输出的正确性,却忽略了中间推理的质量。然而在临床场景中,通过伪造证据或逻辑混乱得出的正确答案与错误答案同样危险。我们提出MedTraj框架,该框架将推理轨迹视为构建、评估和优化的关键对象。该流程从医学推理源生成结构化的多步骤推理链,随后将每条轨迹解析为临床观察、证据、编号推理步骤和最终结论,并从五个质量维度进行评分:连贯性、证据支持、幻觉、完整性和可追溯性。我们通过受控错误注入向原本正确的轨迹中引入针对性故障,以建立特定推理失败与可测量质量下降之间的因果关系。在此基础上,基于边际贡献的步骤级过滤可识别出哪些单个推理步骤会推动或破坏轨迹质量。最后,质量加权上下文学习在推理时将轨迹评估反馈给模型,使其能从优质和劣质推理演示中学习。在CareQA、PubMedQA和CECMed上的实验表明,轨迹上下文可持续提升推理连贯性,与零样本基线相比提升幅度为+0.029至+0.041;在CECMed上,质量加权上下文使正确性几乎翻倍,同时将幻觉比例降低87%。边际贡献分析进一步显示,少数推理步骤承载了大部分质量信号,且将推理链扩展至4步以上会产生收益递减。
英文摘要
Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermediate reasoning. In clinical settings, however, a correct answer reached through fabricated evidence or incoherent logic is as dangerous as an incorrect one. We propose MedTraj, a framework that treats reasoning trajectories as critical objects for construction, evaluation, and optimization. The pipeline generates structured multi-step reasoning chains from medical reasoning sources. Each trajectory is then parsed into clinical observations, evidence, numbered reasoning steps, and a final conclusion, and scored across five quality dimensions: coherence, evidence support, hallucination, completeness, and traceability. Controlled error injection introduces targeted faults into otherwise correct trajectories to establish causal links between specific reasoning failures and measurable quality degradation. Building on this, step-level filtering based on marginal contribution identifies which individual reasoning steps drive or undermine trajectory quality. Finally, quality-weighted context learning feeds trajectory evaluations back into the model at inference time, allowing it to learn from both strong and weak reasoning demonstrations. Experiments across CareQA, PubMedQA, and CECMed demonstrate that trajectory context consistently improves reasoning coherence, with gains of +0.029 to +0.041 over a zero-shot baseline. On CECMed, quality-weighted context nearly doubles the correctness over the zero-shot baseline while cutting the hallucination ratio by 87%. Marginal-contribution analysis further shows that a small minority of reasoning steps carry most of the quality signal, and that extending chains beyond four steps yields diminishing returns.