发表机构
Zhongguancun Laboratory; Tsinghua University; Network Management Center, China Mobile(中关村实验室; 清华大学; 中国移动网络管理中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VRL-Bench基准测试有限试验预算下智能体的试错学习,提出VEX$^2$调度器,在所有六种设置中均优于重试基线。
AI 中文摘要
从试错中学习是提升语言智能体在计算机控制等复杂任务上表现的一种有前景的方法。Reflexion引入了口头强化学习,它将失败的试验转化为文本,以指导后续尝试,而无需更新模型参数。我们提出了VRL-Bench,一个在有限试验预算下公平评估试错学习的测试平台。在MiniWoB和WebShop上,针对三个模型,我们评估了来自多种著名口头记忆方法的更新,这些方法涵盖Reflexion及其后续工作:每种方法在某些设置下相较于无记忆重试提高了观察到的成功率,但在其他设置下则降低了成功率。重放实验表明,使用反思可能会降低成功率,揭示了利用经验与持续探索之间的权衡。我们提出了VEX$^2$,一种口头探索-利用调度器,它使用语言模型联合选择策略并分配剩余的试验预算。VEX$^2$是唯一在所有六种设置中相较于重试实现观察到的成功率正增益的评估更新。
英文摘要
Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.