arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

路由而非修复:面向可靠临床大语言模型答案选择的制度依赖解码校正与轨迹门控路由器

Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection

Zeyu Dong, Benjamin Wang, Joyee W. Jin

arXiv 2609.14825首次发表:更新:

发表机构

Crestwood Preparatory College; St. Theresa of Lisieux Catholic High School; University of Toronto(克雷斯特伍德预备学院; 圣特蕾莎·利雪天主高中; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对临床问答中LLM幻觉问题,提出ALTAS,利用单次前向传播的终端熵和晚期层线性度,按问题在贪心解码与轨迹校正间路由,在保持临床基准不伤害的同时提升真实性。

AI 中文摘要

大型语言模型(LLMs)在临床问答中常因其幻觉倾向而被认为不安全。检索增强、微调和外部验证器需要临床治理机构批准的新基础设施,并可能增加延迟或额外的模型调用。推理时校正利用模型内部的逻辑信号,但固定的变换未必适用于每个问题。一个在真实性压力测试上提高约十个百分点准确率的校正器,在临床多项选择基准上却收益甚微,因为指令微调将输出概率集中在单一答案上,导致终端熵较低。我们提出ALTAS,它通过一次前向传播读取终端熵和晚期层线性度($R^2$),以在贪心解码和晚期层轨迹校正之间为每个问题做出选择。无需训练分类器、探针或头;路由器基于候选答案的逻辑值运行,并增加6.5%的延迟开销。应用于每个问题时,该校正使TruthfulQA在3B和8B规模上分别比贪心解码提高11.4和10.0个百分点($p<10^{-10}$)。按问题门控后,ALTAS保留8.3至9.5个百分点的增益,同时使MedQA、PubMedQA和MedHallu保持在“不伤害”的一个百分点范围内,与贪心解码相比无统计学显著差异。该方法通过了冻结阈值、评分规则和领域标签的验证扫描。

英文摘要

Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation, fine-tuning, and external verifiers require new infrastructure that clinical governance must approve and may add latency or extra model calls. Inference-time correction uses the model's internal logit signals, but a fixed transformation need not suit every question. A corrector that improves accuracy by about ten percentage points on a truthfulness stress test yields negligible gains on clinical multiple-choice benchmarks, where instruction tuning concentrates output probability on one answer and leaves low terminal entropy. We introduce ALTAS, which reads terminal entropy and late-layer linearity ($R^2$) from one forward pass to choose per question between greedy decoding and late-layer trajectory correction. No classifier, probe, or head is trained; the router operates on candidate-answer logits and adds 6.5% latency overhead. Applied to every question, the correction improves TruthfulQA over greedy at 3B and 8B by 11.4 and 10.0 percentage points, respectively ($p<10^{-10}$). Gated per question, ALTAS retains gains of 8.3 to 9.5 percentage points while keeping MedQA, PubMedQA, and MedHallu within a one-percentage-point do-no-harm band, with no statistically significant differences from greedy. The method passes verification sweeps over frozen thresholds, the scoring rule, and the domain label.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑