发表机构
Shenzhen University; The Chinese University of Hong Kong; Dealism(深圳大学; 香港中文大学; 迪利斯姆(Dealism))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出ClinTraceBench基准,评估8种历史表示策略在4种临床大语言模型上的纵向临床推理能力,发现压缩策略存在关系损失与聚合代价,且颠覆了最大主干模型胜出的经验法则。
AI 中文摘要
临床大语言模型助手需对多就诊患者轨迹进行推理,但用于扩展模型的紧凑历史表示(检索、结构化时间线、LLM摘要、智能体记忆)是否保留临床推理所需的纵向信号尚未得到测量。本文提出ClinTraceBench:385个源自MIMIC-IV的带事件ID溯源的验证对话、9项任务分类(T1-T9),以及L0-L4确定性+L5人工审计验证(一致性达98.92%)。在6271个问题(32个单元、200672个预测)上,针对4种主干模型(DeepSeek-V3、GPT-4o-mini、Haiku~4.5、Sonnet~4.6)评估8种历史表示策略:无上下文基准、仅上次就诊、全上下文、BGE-M3密集检索、两种压缩方案、两种智能体记忆系统(Mem0、A-Mem)。得出4项发现:SP4:受控T3注入探针分离出压缩导致的关系损失——在归因语句存在于构建前的情况下,Mem0、A-Mem和llm摘要仅恢复注入正例的0-5.3%;SP1:压缩策略在多就诊趋势和跨患者比较上付出聚合代价;SP2:全上下文的盲态差距达+29.8个百分点(GPT-4o-mini)至+62.7个百分点(Haiku);SP3:弃权(不执行)随上下文长度非单调变化。在Pareto前沿,Haiku在全上下文下优于Sonnet(成本25.76美元vs106.21美元),颠覆了“最大主干模型胜出”的经验法则。
英文摘要
Clinical LLM assistants must reason over multi-visit patient trajectories, yet whether the compact history representations used to scale them---retrieval, structured timelines, LLM summaries, agentic memory---preserve the longitudinal signal clinical reasoning needs has not been measured. We introduce ClinTraceBench: 385 MIMIC-IV-derived verified dialogues with event-ID provenance, a nine-task taxonomy (T1--T9), and L0--L4 deterministic + L5 human-audit validation (98.92\% agreement). We evaluate eight history representation strategies---a no-context floor, \textit{last-visit-only}, \textit{full-context}, BGE-M3 \textit{dense-retrieval}, two compression schemes, and two agentic-memory systems (\textit{Mem0}, \textit{A-Mem})---across four backbones (DeepSeek-V3, GPT-4o-mini, Haiku~4.5, Sonnet~4.6) on 6{,}271 questions: 32 cells, 200{,}672 predictions. Four findings: (SP4) a controlled T3 injection probe isolates compression-induced \textit{relation} loss---with the attribution sentence present \textit{before} construction, \textit{Mem0}, \textit{A-Mem} and \textit{llm-summary} still recover only 0--5.3\% of the injected positives; (SP1) compressed strategies pay an aggregation tax on multi-visit trends and cross-patient comparisons; (SP2) the blind-to-full gap spans $+29.8$~pp (GPT-4o-mini) to $+62.7$~pp (Haiku); (SP3) abstention scales non-monotonically with context length. On the Pareto frontier Haiku dominates Sonnet under \textit{full-context} (\$25.76 vs.\ \$106.21), inverting the ``biggest backbone wins'' heuristic.
CommentsFindings of EMNLP 2026