arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文学习即隐式策略梯度

In-Context Learning as Implicit Policy Gradient

Masahiro Kaneko, Timothy Baldwin

arXiv 2607.23153首次发表:更新:

发表机构

MBZUAI(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型上下文学习现象,证明分数条件上下文学习与策略梯度优化有结构对应,通过自注意力机制等进行分析推导,经多模型实验验证,表明大语言模型能利用分数信息优化输出分布,注意力权重与示例分数相关。

AI 中文摘要

近期研究表明,大语言模型可通过纳入生成样本及其评估分数作为上下文示例来迭代改进输出。但该现象的理论基础尚不清楚。本文表明,分数条件上下文学习与策略梯度优化存在结构对应。首先证明自注意力机制在特定权重矩阵配置下能实现类似REINFORCE算法的奖励加权聚合,并讨论其与预训练变压器行为的关系。这种对应在隐藏状态空间是有方向的,且仅在特定简化条件下成立,我们通过实验量化其强度。在简化隐藏状态模型中,还推导了有界注意力更新引起的分布偏移的精确上界。通过对多个大语言模型的广泛实验验证了理论,表明大语言模型有效利用分数信息将输出分布转向高分示例,且注意力权重与示例分数有强相关性。

英文摘要

Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical foundations underlying this phenomenon remain poorly understood. In this paper, we show that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization. We first provide a constructive proof that self-attention mechanisms can implement reward-weighted aggregation analogous to the REINFORCE algorithm under specific weight matrix configurations, and discuss the relationship between this construction and the behavior of pretrained transformers. The correspondence is directional in hidden-state space and holds exactly only under the stated simplifying conditions; we quantify its strength empirically. Within our simplified hidden-state model, we furthermore derive an exact upper bound on the distribution shift induced by a bounded attention update, yielding a trust-region-like analogy to KL-constrained policy optimization. We validate our theory through extensive experiments across multiple LLMs, demonstrating that LLMs effectively utilize score information to shift output distributions toward high-scoring exemplars, and that attention weights exhibit a strong correlation with example scores.

CommentsCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑