arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习紧凑上下文模型的优化景观

The Optimization Landscape of Learning Compacted Context Models

Thomas Villeneuve, Alex Sandomirsky, Charles O'Neill, Max Kirkby, Michael Psenka

arXiv 2610.05885首次发表:更新:

发表机构

Base Labs(Base Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文分析持续学习中KV缓存压缩优化问题的困难性,提出简化Perceiver架构,在多个领域任务上匹配或超越基线性能。

AI 中文摘要

许多工作通过无限上下文窗口的视角来处理持续学习。当智能体将更多观察放入上下文(具体而言是KV缓存)时,压缩该上下文类似于直接进行记忆操作,而不影响基础模型的权重。许多工作将KV压缩视为一个优化问题:学习一组较小的KV向量,使其匹配完整KV缓存的行为。虽然这保留了基础模型的行为,但通过冻结的基础模型进行优化会导致一个高度非平凡的优化问题,其损失景观脆弱且平坦。在本文中,我们刻画了使这些优化问题困难的因素,并证明了一个高度简化的基于Perceiver的架构不仅在连续上下文压缩中匹配完整Perceiver变压器的性能,而且在压缩效用上优于基线。结果在金融、法律、古腾堡和代码领域的多项选择题任务上呈现。

英文摘要

Many works approach continual learning through the lens of infinite context windows. As an agent puts more observation into context (concretely the KV cache), compacting said context is akin to direct memory manipulation, without affecting the base model's weights. Many works pose KV compaction as an optimization problem: learn a smaller set of KV vectors that matches the behavior of the full KV cache. While this preserves base model behavior, optimizing through a frozen base model results in a highly nontrivial optimization problem with a brittle and flat loss landscape. In this paper, we characterize what makes these optimization problems difficult and demonstrate that a heavily simplified Perceiver-based architecture not only matches performance of a full Perceiver transformer in continuous context compaction, but outperforms baselines on compaction utility. Results are presented on MCQ tasks across Finance, Legal, Gutenberg, and Code.

CommentsNeurIPS 2026 Workshop: Continual Learning in the Era of Foundation Models and Embodied Agents

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑