arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时适配:用于保留性能的领域专业化的条件记忆适配器

When to Adapt: Conditional Memory Adapters for Retention-Preserving Domain Specialization

Jiayu Hou, Lei Wang

arXiv 2608.29327首次发表:更新:

发表机构

University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型参数高效微调导致域外性能下降的问题,提出Engram Adapter框架,通过条件激活实现领域专业化,在提升领域内准确率的同时保留绝大多数域外性能。

AI 中文摘要

部署在专业领域的大语言模型必须提升领域内性能,同时不牺牲通用能力。现有的参数高效微调方法通常始终启用:其学习到的扰动会应用于所有输入,这可能会降低域外(OOD)性能。我们提出Engram Adapter,这是一种将预训练时的条件记忆重新用作冻结大语言模型(LLM)的事后适配器的框架。它对局部n元语法模式使用多通道匹配,并结合显式占用跟踪作为轻量级选择性先验,使得残差注入更可能发生在领域内输入上,同时学习到的标量门抑制不连贯的域外检索。我们在Qwen3-4B和Qwen3-8B上进行评估,以AG-News和MedMCQA作为适配任务,域外基准涵盖推理、翻译、代码生成和法律推理。Engram Adapter提升了领域内准确率,同时保留了99.4%至100.1%的平均域外性能;在LegalBench上,它的平均表现略优于冻结的基础模型,而可比的始终启用基线则大幅下降。机制分析表明,尽管域外激活非零,但门和投影衰减将残差降低到隐藏状态范数的约0.08%,产生小的KL漂移和可忽略的准确率变化。这些结果表明,条件激活是在冻结主干上实现模块化、保留性能的领域专业化的有前景的途径。

英文摘要

Large language models deployed in specialized domains must improve in-domain performance without sacrificing general capabilities. Existing parameter-efficient fine-tuning methods are typically always on: their learned perturbations are applied to every input, which can degrade out-of-domain (OOD) performance. We propose Engram Adapter, a framework that repurposes pretraining-time conditional memory as a post-hoc adapter for frozen LLMs. It uses multi-channel matching over local n-gram patterns with explicit occupancy tracking as a lightweight selectivity prior, making residual injection more likely on in-domain inputs while a learned scalar gate suppresses incoherent OOD retrievals. We evaluate on Qwen3-4B and Qwen3-8B with AG-News and MedMCQA as adaptation tasks and OOD benchmarks spanning reasoning, translation, code generation, and legal reasoning. Engram Adapter improves in-domain accuracy while preserving 99.4%--100.1% of average OOD performance; on LegalBench it slightly exceeds the frozen base model on average, whereas comparable always-on baselines degrade sharply. Mechanistic analyses show that although OOD activations are non-zero, gate and projection attenuation reduce residuals to approximately 0.08% of hidden-state norm, yielding small KL drift and negligible accuracy change. These results suggest conditional activation is a promising route toward modular, retention-preserving domain specialization over frozen backbones.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑