arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码代理是强大的提示优化器

Coding Agents are Strong Prompt Optimizers

Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani

arXiv 2609.26261首次发表:更新:

AI 中文总结

本研究提出编码代理技能蒸馏(CASD),通过静态轨迹语料库的统计反思直接合成优化提示,无需环境访问或验证数据,在多个基准上优于现有方法,且成本大幅降低。

AI 中文摘要

基于搜索的提示优化器通过迭代搜索来改进提示:它们提出编辑,执行新的轨迹,对产生的轨迹进行评分,并仅保留能改善验证指标的编辑。我们表明,这种优化循环是不必要的。仅给定静态的代理轨迹语料库,一个现成的编码代理可以直接合成优化后的提示,既不需要环境访问,也不需要验证数据。我们将这种方法称为编码代理技能蒸馏(CASD)。关键见解在于反思范围。编码代理不是在每个优化步骤中对一小批轨迹进行推理,而是编写并执行分析代码以计算语料库范围的统计信息,识别系统性故障模式,检查代表性情节,并将所得见解提炼为行为规则。在四个智能体基准测试(ALFWorld、τ²-bench零售和电信以及SpreadsheetBench-Verified)中,在匹配的数据访问条件下,单次CASD通过在四个基准中的三个上优于最先进的反思性提示优化器GEPA,并在所有四个基准上优于验证门控的反思性搜索(SkillOpt),将未优化的基线平均提高16.6个百分点,而GEPA为10.9,SkillOpt为5.3。由于CASD执行单次离线分析而不是迭代搜索,生成优化提示的成本约为1.60美元,比验证门控搜索便宜22倍以上。即使当竞争方法被授予额外的验证数据和无限制的环境访问时,CASD在四个基准中的两个上仍然领先。这些结果表明,语料库规模的统计反思是迭代搜索进行提示优化的可行替代方案。

英文摘要

Search-based prompt optimizers improve prompts through iterative search: they propose edits, execute fresh rollouts, score the resulting trajectories, and retain only edits that improve a validation metric. We show that this optimization loop is unnecessary. Given only a static corpus of agent trajectories, an off-the-shelf coding agent can directly synthesize an optimized prompt, requiring neither environment access nor validation data. We call this approach \textit{Coding-Agent Skill Distillation} (CASD). The key insight is reflection scope. Rather than reasoning over a small batch of trajectories at each optimization step, the coding agent writes and executes analysis code to compute corpus-wide statistics, identifies systematic failure modes, inspects representative episodes, and distills the resulting insights into behavioral rules. Across four agentic benchmarks (ALFWorld, $τ^2$-bench retail and telecom, and SpreadsheetBench-Verified), under matched data access, a single CASD pass outperforms GEPA, a state-of-the-art reflective prompt optimizer, on three of four benchmarks and outperforms validation-gated reflective search (SkillOpt) on all four, improving the unoptimized baseline by 16.6 percentage points on average versus 10.9 for GEPA and 5.3 for SkillOpt. Because CASD performs a single offline analysis pass rather than iterative search, producing an optimized prompt costs approximately \$1.60---over $22\times$ cheaper than validation-gated search. Even when competing methods are granted additional validation data and unrestricted environment access, CASD remains ahead on two of four benchmarks. These results suggest that corpus-scale statistical reflection is a viable alternative to iterative search for prompt optimization.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑