arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型能否在长时程上进行推理?纵向临床推理中上下文策略的实证评估

Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning

Taye Akinrele, Noorbakhsh Amiri Golilarz, Subash Neupane, Sudip Mittal, Shahram Rahimi

arXiv 2610.00562首次发表:更新:

发表机构

The University of Alabama(阿拉巴马大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在MedLoCoMo上比较五种上下文策略,发现情节性与混合策略在纵向临床推理中准确率最高,且证据选择与呈现比单纯增加上下文长度更关键。

AI 中文摘要

纵向临床推理要求大语言模型(LLMs)识别并整合分布在扩展患者病史中的相关证据。尽管长上下文模型能够处理日益增多的信息,但提供更多病史并不必然使相关证据更易获取或改善推理。我们在MedLoCoMo上比较了五种上下文策略(完整、近期、情节性、语义和混合),涉及四个开放权重的大语言模型,考察了答案正确性、对查询-证据距离的鲁棒性,以及在对无支撑前提的问题上的弃权(不执行)表现。情节性和混合策略通常达到最强的整体准确率,而随着支持性证据距离增大,近期上下文策略的性能下降最为严重;情节性和混合策略在长距离上保持最高准确率。对抗性问题的分析进一步表明,在可回答问题上的强表现并不必然转化为在可用病史不支持所请求结论时成功弃权(不执行)。这些发现表明,可靠的纵向推理不仅取决于大语言模型能访问多少病史,更关键地取决于相关证据如何被选择和呈现以用于推理。

英文摘要

Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑