arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少样本演示激发大语言模型中上下文世界表征的使用

Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs

Kohsei Matsutani, Gouki Minegishi, Core Francisco Park, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

arXiv 2609.24352首次发表:更新:

发表机构

The University of Tokyo; Prior Computers; Harvard University(东京大学; Prior Computers; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现少样本演示能提升LLM在图追踪等任务中的上下文世界表征使用,通过线性探测揭示其正交移动机制,并验证了在ARC-AGI等任务上的泛化效果。

AI 中文摘要

大语言模型(LLMs)在作为智能体时,被期望利用上下文中的观测数据,推断世界背后的潜在状态空间,并利用其进行下游预测。然而,先前的研究表明,在图追踪任务中,LLMs难以使用在上下文中学习到的表征,该任务要求模型构建控制数据生成过程的图的表征,并将其用于后续预测。在本文中,我们表明,将此扩展到少样本设置(其中每个演示由具有相同或不同图拓扑的不同世界生成)可提升来自4个模型家族的6个模型的预测性能。为了理解这一改进,我们线性探测了一个在隐藏状态中编码图信息的低维世界表征。值得注意的是,我们发现少样本演示重新定位了世界表征并增加了其预测用途。具体而言,对于每个模型,这些世界表征的移动方向几乎与其原始子空间正交,且对这些表征的干预比对其他子空间的干预更显著地选择性损害性能。与此见解一致,我们表明,具有来自不同世界的观测的少样本演示可提升ARC-AGI-1&2、Web智能体任务和Othello上的性能。我们的发现阐明了少样本演示在上下文世界建模中的作用和内部机制。更广泛地说,我们的工作推进了对LLM智能体如何从上下文观测中学习的理解,并为其进一步改进提供了启示。

英文摘要

Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction. However, prior work demonstrated that LLMs struggle to use representations learned in context on a graph tracking task, where the model needs to construct a representation of the graph governing data generation process and use it for subsequent predictions. In this paper, we show that extending this to few-shot settings, where each demonstration is generated from a different world with either the same or different graph topologies, enhances its prediction on 6 models from 4 model families. To understand this improvement, we linearly probe a low-dimensional world representation that encodes graph information in the hidden states. Notably, we find that few-shot demonstrations relocate the world representation and increase its predictive use. Specifically, for each model, these world representations shift in directions nearly orthogonal to their original subspace, and interventions on these representations selectively impair performance more than interventions on other subspaces. Consistent with this insight, we show that few-shot demonstrations with observations from different worlds improve performance on ARC-AGI-1&2, web agent tasks, and Othello. Our findings elucidate the role and internal mechanisms of few-shot demonstrations in in-context world modeling. More broadly, our work advances our understanding of how LLM agents learn from in-context observations and provides implications for their further improvement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑