发表机构
Renmin University of China; Tencent(中国人民大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体部署中测试时学习延迟与不可靠复用问题,提出非参数框架StepLearn,分离即时使用与持久信任,通过前瞻验证更新外部知识,在WebArena-Lite和ALFWorld上显著超越基线。
AI 中文摘要
在部署期间适应大语言模型智能体,不仅需要保留过去的经验,还需要将新的观察转化为及时的指导。然而,许多测试时学习方法从已完成的回合中获取知识。因此,来自正在进行交互的反馈可能无法及时提炼为知识,以帮助下一步决策。在单个转移的粒度上获取知识可以减少这种延迟,但提出了另一个挑战:在一个回合中有用的规则可能不足以可靠地指导未来的回合。等待验证可能会丧失即时收益,而不受限制的复用则可能传播偶然或错误归因的指导。我们引入了StepLearn,一个非参数框架,将即时使用与持久信任分离。它将信息丰富的转移转化为可以指导下一步的假设,同时要求在跨回合复用之前进行前瞻性验证。其预测效果与源回合之外的后续观察进行核对,只有得到充分支持的假设才能用于持久指导。这一过程在保持所有模型参数固定的同时更新外部知识。在WebArena-Lite和ALFWorld上的五轮实验中,StepLearn分别使用GPT-5-mini和Qwen3.5-35B-A3B实现了59.9%和84.0%以及57.8%和88.1%的平均成功率。在四种设置中,它比最强基线EvoTest高出2.2-12.7个百分点。学习动态进一步表明,这些收益不仅限于最终重复,在大多数设置的首次任务尝试中就已经存在优势。
英文摘要
Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.