arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向临床诊断的大语言模型训练中知识图谱的外科式对齐

Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models

Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao

arXiv 2608.26587首次发表:更新:

发表机构

University of Colorado Anschutz Medical Campus; University of Wisconsin Madison; Loyola University(科罗拉多大学安舒茨医学校区; 威斯康星大学麦迪逊分校; 洛约拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对临床诊断的LLM训练,通过多维度系统研究引入GID和GD,发现KL正则化下KG判断训练的稀疏更新(外科式对齐)可提升推理质量,需结合优化几何诊断评估KG-LLM整合。

AI 中文摘要

生物医学知识图谱(KG)提供结构化的医学知识,可在临床诊断应用中为大语言模型(LLM)的推理提供依据,但如何将KG信号整合到LLM中仍是一个开放问题。我们开展了一项系统研究,涵盖5种KG任务形式、3种训练范式、2个KG和3个基础LLM。在任务层面,所有范式均优于未微调的基线,但具有可比域内准确率的方法表现出显著不同的知识迁移行为。我们引入梯度干预密度(GID)和梯度畸变(GD),以衡量优化器对预训练模型的修改范围。GID和GD共同揭示了清晰的划分:在KL正则化下的KG判断训练会产生稀疏、局部的更新(我们将这种机制称为外科式对齐),而特定任务的SFT则产生密集的更新。受控消融实验表明,目标函数和KL项对稀疏性的贡献是独立的,且产生稀疏更新的范式即使在域内准确率低于特定任务SFT时,也能提升推理质量。因此,评估KG-LLM整合需要将准确率与优化几何诊断相结合。我们的实现可在该https URL获取。

英文摘要

Biomedical knowledge graphs (KGs) offer structured medical knowledge that can ground large language model (LLM) reasoning in clinical diagnosis application, yet how KG signal should be integrated into LLMs remains an open question. We present a systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs. At the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior. We introduce Gradient Intervention Density (GID) and Gradient Distortion (GD) to measure how broadly an optimizer modifies the pretrained model. GID and GD together reveal a clear divide: KG-judgment training under KL regularization produces sparse, localized updates (a regime we term as surgical alignment), while task-specific SFT produces dense ones. A controlled ablation shows that the objective and KL contribute to sparsity independently, and the paradigms that produce sparse updates also improve reasoning quality, even when their in-domain accuracy is lower than task-specific SFT. Assessing KG-LLM integration thus requires complementing accuracy with optimization-geometry diagnostics. Our implementation can be found at https://github.com/LARK-NLP-Lab/Surgical-Alignment.

CommentsThis work has been accepted to EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑