arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向下游的上下文选择用于在线上下文强化学习

Downstream-Aware Context Selection for Online In-Context Reinforcement Learning

Ruihan A. Li, Shangtong Zhang, Rohan Chandra

arXiv 2609.33166首次发表:更新:

AI 中文总结

提出有界历史上下文管理框架,预测移除历史交互的下游效应,动态决定保留历史量,在驾驶和ScienceWorld任务中减少令牌使用而不损性能。

AI 中文摘要

上下文强化学习(ICRL)使大型语言模型代理能够利用其交互历史适应新环境,而无需更新模型参数。然而,反复基于不断增长的历史进行条件化处理会导致大量的令牌成本。我们提出了一种有界历史上下文管理框架,该框架预测移除历史交互的任务相关下游效应,以指导历史选择并确定决策相关的上下文预算。形式上,我们的框架使用完整的滚动历史作为参考。预测器评估移除效应,定义删除顺序,并应用共享的选择标准来确定每个决策时应保留多少历史。我们在封闭循环SUMO驾驶环境中,在保留分布、未见领域和未见路线设置下评估该方法,并在ScienceWorld中采用持续ICRL协议进行评估。相对于使用完整上下文的基线,我们的方法在三种驾驶设置中将总令牌使用量分别减少了25.7%、25.8%和23.2%,同时保持了相当的封闭循环驾驶性能。在ScienceWorld中,与完整上下文相比,它将总令牌使用量减少了52.1%,并且分别比Recent和Similarity基线少使用30.2%和37.8%的令牌,同时保持了性能。

英文摘要

In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework that predicts the task-dependent downstream effect of removing historical interactions to guide history selection and determine a decision-dependent context budget. Formally, our framework uses the full rolling history as a reference. The predictor evaluates removal effects, defines a deletion ordering, and applies a shared selection criterion to determine how much history to retain at each decision. We evaluate the method in closed-loop SUMO driving under held-out in-distribution, unseen-domain, and unseen-route settings, and in ScienceWorld under a continual ICRL protocol. Relative to a baseline using the full context, our method reduces total token usage by 25.7%, 25.8%, and 23.2% across the three driving settings while maintaining comparable closed-loop driving performance. In ScienceWorld, it reduces total token usage by 52.1% compared to full context and uses 30.2% and 37.8% fewer tokens than the Recent and Similarity baselines, respectively, while maintaining performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑