arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动构建文本操纵器:通过代码自博弈为黑盒优化提炼文本操纵器

Code-to-Harness: Distilling Black-Box Optimizers from Self-Play

Yi Wu, Zheng Ren, Zhiyu Hu, Haochen Wang, Daryl Chang, Li Wei, Ting Wang, Zhen Li, Pooja Gupta, Nitin Jindal, Lukasz Heldt

arXiv 2609.09468首次发表:更新:

发表机构

Google Mountain View, USA(谷歌山景城)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过代码自博弈实践提炼文本操纵器,将搜索策略以文本形式迁移至不同模型,显著降低黑盒优化遗憾值。

AI 中文摘要

一个智能体能否通过可执行的实践学习数值搜索策略,然后将该策略以文本形式迁移?我们研究低预算黑盒优化,在此场景下,未经辅助的语言模型仍远低于经典优化器的性能。在开发过程中,智能体反复编写并评估优化器程序。随后,它将该程序及实践记录一次性提炼为一份197词的初级操纵器A,该操纵器在评估前被冻结。在独立的N=30研究中,操纵器A将Gemini Flash的遗憾值降低了48%(p<.001),在实践族上进入GP-BO性能范围,并在所有三个保留的BBOB景观上降低了平均遗憾值。同一文本改善了所有测试的Gemini执行器,并迁移至Claude Sonnet,遗憾值分别降低43%和49%(p≤.005)。一项独立的端到端复现产生了操纵器B,这是同一性能层级上的不同程序与文本。同一框架还在一个密封的YouTube奖励调优生产基准上取得了最低遗憾值。因此,可执行实践是发现搜索策略的可行途径,而语言是部署该策略的可移植媒介。

英文摘要

Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It then distills the resulting program and practice record once into a 197-word primary Harness A, which is frozen before evaluation. Harness A reduces Gemini Flash regret by 48\% in an independent $N=30$ study ($p<.001$), enters the GP-BO performance range on the practice family, and lowers mean regret on all three held-out BBOB landscapes. The same text improves every tested Gemini executor and transfers to Claude Sonnet, reducing regret by 43\% and 49\% ($p\leq.005$). An independent end-to-end replication produces Harness B, a different program and text at the same performance tier. The same framework also attains the lowest regret on a sealed YouTube reward-tuning production benchmark. Executable practice is thus a viable way to discover a search policy, and language a portable medium for deploying it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑