MIITA:用于小语言模型持续学习的内存诱导推理时间自适应
MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
- Baylor University(贝勒大学)
- NEC Laboratories America(美国 NEC 实验室)
- University of Arkansas(阿肯色大学)
- Southern Illinois University(南伊利诺伊大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对小语言模型持续学习中参数更新致灾难性遗忘及内存不足问题,提出MIITA框架,它通过存储校正方向原型并在推理时检索应用,结合理论分析,在多种设置实验中提升性能并减轻遗忘。
AI中文摘要:
持续学习(CL)对于小语言模型(SLM)在资源受限部署中适应不断变化的现实世界需求至关重要。然而,直接更新其有限的参数空间会导致灾难性遗忘。基于内存的方法通过将知识保留与参数解耦自然地解决了这个问题,但现有的针对大语言模型(LLM)设计的方法依赖于丰富的存储和强大的上下文推理,而这是SLM所缺乏的。为应对这些挑战,我们提出了MIITA,一种用于受限存储下监督式CL的内存诱导推理时间自适应框架。MIITA将监督经验存储为带有语义锚点的紧凑校正方向原型,并在推理时使用基于语义和不确定性的线索进行检索。检索到的方向通过门控临时隐藏状态自适应应用,无需更新主干、扩展提示或测试时反向传播即可无损重用过去的监督。局部理论分析将此设计与一阶损失减少、不确定性引导检索和保留旧阶段知识的方向覆盖联系起来。在各种监督式CL设置上的广泛实验表明,MIITA在固定内存预算下持续提高最终性能并减轻遗忘。
英文摘要:
Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models (LLMs) rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.