发表机构
University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出任务算子(TO),通过将注意力头输出视为上下文掩码对应物的仿射变换并解析重放,压缩上下文学习,在词汇、算法和推理任务上优于先前方法,并缩小了与零样本推理的差距。
AI 中文摘要
上下文学习(ICL)使语言模型能够从演示中执行新任务而无需更新权重。然而,每次ICL推理都需要处理完整的示例集,导致部署效率低下,并且ICL的机制尚未被完全理解。先前的工作将ICL压缩为从特定层或位置提取的固定激活向量,但这些与输入无关的干预在输出依赖于与输入的细粒度交互的复杂任务上失败。通过分析ICL前向传播,我们表明每个注意力头的输出是其上下文掩码对应物的仿射变换,并且该变换的参数在给定任务的样本间经验上是稳定的。基于此,我们引入了任务算子(TO),它将此变换作为解析推导的更新重放到注意力输出投影中。在词汇、算法和推理任务中,TO在先前方法中取得了最佳整体性能,并大幅缩小了零样本推理与ICL之间的差距。我们进一步表明,提取的知识集中在跨层和位置的任务特定稀疏电路中,并且对来自不相交演示批次的算子进行平均,可以在不扩展上下文窗口的情况下实现有效的多示例缩放。我们的代码可在该URL获取。
英文摘要
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.
CommentsNeurIPS 2026