arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34126cs.SEcs.AI

JET:测试时基于法官引导的智能体程序进化

JET: Judge-Guided Evolution at Test Time for Agent Programs

Yao Long Teng, Jiayi Cai, Bo An

首次发表
浏览论文内容

中文总结 AI 辅助

JET通过训练并冻结可执行法官,在测试时引导智能体程序进化,无需目标评估器,在WebShop任务上显著提升奖励与成功率。

中文摘要 AI 辅助

智能体的可执行程序决定了它如何使用工具、处理观察结果以及应对失败。在测试时进化该程序有助于适应,但在无法获得真实奖励的情况下,决定保留哪些更改是困难的。执行轨迹提供了智能体行为的证据,但解释这些证据需要一个在任务和候选程序变化时仍然有用的法官。我们提出了测试时法官引导进化(JET),该方法在标记的源轨迹上进化一个可执行法官,然后冻结并转移它以指导目标端的程序进化。法官提供评分和诊断反馈,无需访问目标评估器或更新模型权重。在未见过的WebShop任务上,当进化从未进化程序(冷启动)开始时,JET的平均奖励比固定标准指导高出约13%,当从已在源任务上优化的程序(热启动)开始时,高出4%,冷启动精确成功率的相对改进为36%。在PushT上,一个精确法官对照(法官从观察中重建评分规则)表明,在没有法官错误的情况下,程序搜索成为瓶颈。分析识别了进化代码中有用的奖励预测逻辑,并表明仅靠更好的最终选择无法解释这些收益。这些结果支持在保持评估器的任务转移下,可执行法官转移用于程序适应。

英文摘要

An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior, yet interpreting that evidence requires a judge that remains useful as tasks and candidate programs change. We introduce Judge-Guided Evolution at Test Time (JET), which evolves an executable judge on labeled source trajectories, then freezes and transfers it to guide target-side program evolution. The judge supplies scores and diagnostic feedback without target evaluator access or model-weight updates. On unseen WebShop tasks, JET achieves approximately 13% higher mean reward than fixed-rubric guidance when evolution begins from an unevolved program (cold start) and 4% higher when it begins from one already optimized on source tasks (warm start), with a 36% relative improvement in cold-start exact success. An exact-judge control on PushT, where the judge reconstructs the scoring rule from observations, shows that without judge error, program search becomes the bottleneck. Analyses identify useful reward-prediction logic in the evolved code and show that better final selection alone cannot explain the gains. These results support executable judge transfer for program adaptation under evaluator-preserving task shifts.

发表机构

  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑