发表机构
Delhi Technological University(德里理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究训练小型网络将上下文学习摊销为潜在任务表征,发现生产性语言规则(如屈折、词形还原)可有效摊销甚至超越ICL,而任意配对(反义)无法摊销。
AI 中文摘要
上下文学习(ICL)可以被摊销为潜在对象(任务向量、函数向量、上下文向量),以零样本推理成本恢复少样本行为,但近期理论表明,静态向量仅相当于一条合成示范,在高秩映射(如词级双射)上必然失败。我们提出该问题的语言学版本:哪些语言操作可以从提示中摊销出来?我们训练了一个260万参数的网络,它读取少样本支持集的几何结构(质心、主子空间、谱,计算一次并缓存),并在冻结的GPT-2-large/XL的中层深度,对查询的残差流产生输入条件化的加性更新。在八个屈折方向和一种词汇关系上,在阻止方向间反转对泄漏的标准划分下,出现了三种机制。在前向屈折上,10-shot ICL较强(0.67-0.89),而提取的任务向量崩溃(≤0.06),该变换以严格零样本的每查询成本匹配ICL。在词形还原方向上,冻结的GPT-2可以执行,但10个示范系统性地无法传达(1.5B规模下ICL为0.13-0.48),该变换完全不受ICL上限约束:它达到0.78-0.92,比ICL高出多达72个百分点(过去到现在:0.85对0.13)。在任意配对(反义关系)上,每个摊销器在所有规模、容量和种子下都稳定在ICL的一半附近。控制实验表明,支持流形作为因果必要的任务指纹:错误任务流形将准确率降至≤0.06,仅查询的变体无法区分共享输入空间的任务,留一任务外转移为零。生产性规则摊销为潜在任务表征,有时比提示能传达的更好;记忆化的配对则不能。
英文摘要
In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections. We ask a linguistic version of this question: which linguistic operations can be amortized out of the prompt? We train a 2.6M-parameter network that reads the geometry of a few-shot support set (centroid, principal subspace, spectrum, computed once and cached) and produces an input-conditioned additive update to the query's residual stream at a mid-depth layer of a frozen GPT-2-large/XL. Across eight inflectional directions and one lexical relation, under a canonical split that bars inverted-pair leakage between directions, three regimes emerge. On forward inflection, where 10-shot ICL is strong (0.67-0.89) and extracted task vectors collapse (<=0.06), the transform matches ICL at strictly zero-shot per-query cost. On lemmatization directions, which frozen GPT-2 can execute but 10 demonstrations systematically fail to convey (ICL 0.13-0.48 at 1.5B), the transform is not capped by ICL at all: it reaches 0.78-0.92, up to +72 points over ICL (past to present: 0.85 vs. 0.13). On arbitrary pairings (antonymy) every amortizer plateaus near half of ICL at every scale, capacity, and seed tested. Controls show the support manifold acts as a causally necessary task fingerprint: wrong-task manifolds collapse accuracy to <=0.06, query-only variants cannot disambiguate tasks sharing an input space, and leave-one-task-out transfer is zero. Productive rules amortize into latent task representations, sometimes better than prompting can convey them; memorized pairings do not.
CommentsAccepted to the NeurIPS 2026 Workshop on Linguistic Principles for Foundation Models (LP4FM). 5 pages