AI 中文总结
本研究提出DSSR方法,通过前向滚动评分训练写入者减少写入时遗憾,在TextWorld任务中提升智能体性能,并揭示了信用分配对延迟场景的限制。
AI 中文摘要
长任务产生的历史信息超过了LLM智能体上下文能容纳的量,也超过了即使历史信息能容纳时智能体可靠使用的量。因此,越来越多的研究工作让智能体携带一个简短的书面状态:每一步,一个写入者重写状态,而一个读取者仅根据该状态行动。步骤保持廉价,但写入者丢弃的任何信息都会在后续决策揭示其需要之前丢失。我们量化了这种损失,并询问训练是否能减少它。将写入的状态与事后以相同大小写入的最佳状态进行比较,我们将读取者的损失分为预算损失(任何该大小的状态都必须承担的损失)和写入时遗憾(源于写入者选择的损失)。在我们控制一个事实在被需要之前必须携带多长时间的TextWorld烹饪游戏中,一个持有这些事实的128个令牌的状态几乎赢得了所有游戏,而提示式语言模型写入者最多赢得17%的游戏。几乎所有的损失都是写入时遗憾,并且它随着延迟的增长而增加。然后,我们根据读取者自身的损失来训练写入者。DSSR(决策充分的状态表示)根据读取者在写入者将候选状态携带向前后的表现来对候选状态进行评分,并教导写入者偏好更好的状态。这种前向滚动评分预测了游戏结果(ρ=0.48),而像事后方法通常所做的那样,将候选作为固定上下文进行评分则不能预测(ρ≤0.07)。在一个预先注册的测试分割上(仅打开一次),当事实很快需要时,训练增加了+7.0 [+1.9, +12.2]个百分点的成功率,使一个简单的摘要写入者达到了基于信念和槽位的内存提示的水平。随着延迟的增长,这种增益缩小,并且仅在最短延迟下显著。我们将这一限制追溯到信用分配:现在保留一个事实只有在每次后续重写都保留它时才有回报,而单步评分无法看到这一点。
英文摘要
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader's loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer's choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader's own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes ($ρ= 0.48$), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not ($ρ\leq 0.07$). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.
Comments29 pages (10 main, 17 appendix), 12 figures (4 main, 8 appendix), 18 tables (2 main, 16 appendix)