arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究线性探测对医学问答中语言语域、医学专科及语料库迁移的鲁棒性

Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto

arXiv 2609.01361首次发表:更新:

发表机构

Amsterdam UMC, University of Amsterdam; Amsterdam Public Health(阿姆斯特丹大学医学中心、阿姆斯特丹大学; 阿姆斯特丹公共卫生研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究线性探测在医学问答中对语域、医学专科及语料库迁移的鲁棒性,发现真值方向在医学领域内大体稳定,但部分语料库迁移会导致其性能下降,信号部分与数据集结构绑定。

AI 中文摘要

基于大型语言模型(LLM)隐藏状态训练的线性分类器(线性探测)可通过单次前向传播标记事实错误。从几何角度看,这意味着真假陈述在隐藏状态空间中沿稳定方向(即“真值方向”)分离。现有研究对该方向是否能跨输入迁移存在分歧,但因跨数据集探测迁移实验同时混淆了多种输入变化,难以解读分歧原因。本研究在医学问答(QA)中分离出三类变量:写作风格(语域)、领域(医学专科)和语料库(数据集)。我们构建了一个基准,包含500条MedQA条目,每条重写为四种风格(教科书、患者、临床笔记、口语),标注医学专科,并与另外两个考试语料库MedMCQA和MMLU-medical分组用于跨数据集评估。对四个开放权重LLM(20亿至80亿参数)进行探测后发现,真值方向对写作风格鲁棒性较强(保留事实的平均Δ_语域≈0.10 AUROC),对医学专科也较鲁棒(Δ_专科≈0.03),但在不同语料库上表现不均:在MMLU-medical上AUROC下降0.12,在MedMCQA上下降0.21,约为语域差异的两倍。该语域结果在第二个生成器上可复现,且适用于人类撰写的患者问题。因此,真值方向在医学领域内大体稳定,但在部分语料库迁移下会失效,且问题格式无法解释该失效,这表明线性探测恢复的信号部分与数据集结构绑定,而非仅依赖医学知识。

英文摘要

Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. Geometrically, that implies that true and false statements separate along a stable direction in hidden state space, i.e., the truth direction. Prior work disagrees on whether this generalises across input shifts, but the disagreement is hard to interpret because cross-dataset probe transfer experiments confound several kinds of input change at once. We isolate three such variables in medical question-answering (QA): writing style (register), domain (medical specialty), and corpus (dataset). We build a benchmark using 500 MedQA entries, each rewritten into four styles (textbook, patient, clinical note, colloquial), annotated with clinical specialty, and grouped with two other exam corpora, MedMCQA and MMLU-medical, for cross-dataset evaluation. Probing four open-weight LLMs (2--8B), we find that the truth direction is largely robust to writing style (mean $Δ_\text{register} \approx 0.10$ AUROC on held-out facts) and to medical specialty ($Δ_\text{specialty} \approx 0.03$), but degrades unevenly across corpora: by $0.12$ AUROC on MMLU-medical and by $0.21$ on MedMCQA, roughly twice the register gap. The register result replicates with a second generator and carries over to human-written patient questions. The truth direction is therefore largely stable within the medical domain but breaks under some corpus shifts, and question format does not explain the break, which suggests that the signal a linear probe recovers is partly bound to dataset structure rather than to medical knowledge alone.

CommentsAccepted to EMNLP 2026 (Main Conference). 9 pages, 4 figures (main text); 28 pages total, including appendix. Code and data: https://github.com/mnishant2/MedProbe_release

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑