大型语言模型中的领域特定术语:通用模型与专业模型的比较分析
Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models
- University of Chicago(芝加哥大学)
- Toyota Technological Institute at Chicago(芝加哥丰田技术学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究通过两个医学术语基准比较通用与医学微调Llama-3.1模型,发现微调未提升术语理解,机制分析揭示组件失衡,重加权可弥补差距,表明领域适应不保证专业术语性能提升。
中文摘要 AI 辅助
大型语言模型(LLMs)在通用任务上表现出卓越的能力,然而它们在高度专业化的技术领域中的性能往往会下降。此外,关于领域特定术语的参数化知识如何在这些模型中被编码,我们知之甚少。我们通过贡献两个新颖的医学术语评估基准来填补这一空白,并将一个通用型Llama-3.1模型与一个在医学领域数据上微调的变体进行了评估比较。令人惊讶的是,通用模型在这两项任务上的表现均优于医学微调模型。利用机制可解释性工具,我们发现了医学微调模型存在系统性校准错误的模式。微调模型并未重组参数化知识,而是将更大的权重放在与术语偏好预测相关的一小部分模型组件上。我们发现,针对基准任务应用组件重新加权策略能够成功抑制这些组件,并缩小与通用基线的差距。我们还观察到,一些对术语敏感的组件将知识迁移到涉及材料科学术语的相同任务中,这表明它们编码了一种部分领域无关的专业术语概念。我们的结果提供了一个案例研究,其中医学微调的检查点在术语理解方面并未优于其通用对应模型,这凸显了领域适应不应被假定为能在专业术语上带来更好的性能。
英文摘要
Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within these models. We address this gap by contributing two novel medical jargon evaluation benchmarks and evaluate a general-purpose Llama-3.1 model against a variant fine-tuned on medical-domain data. Surprisingly, the general-purpose model outperforms the medically fine-tuned model on both tasks. Using mechanistic interpretability tools, we find systematic patterns of miscalibration for the medically fine-tuned model. Instead of reorganizing parametric knowledge, the fine-tuned model places greater emphasis on a small subset of model components associated with jargon-favoring predictions. We find that applying component reweighting strategies against the benchmark tasks successfully suppresses these components and closes the gap with the general-purpose baseline. We also observe that some jargon-sensitive components transfer knowledge to the same tasks involving materials science jargon, suggesting they encode a partially domain-agnostic notion of specialized terminology. Our results provide a case study in which a medically fine-tuned checkpoint does not improve jargon comprehension over its general-purpose counterpart, highlighting that domain adaptation should not be assumed to yield better performance on specialized terminology.