arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06714cs.AI

优化器即智能体:跨提示、程序与机器学习工作流的推理驱动搜索

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ReASearch统一框架,让智能体自主优化提示、程序与ML工作流,在14项任务中优于专用系统,部分发现超人类最佳结果的方案。

中文摘要 AI 辅助

当前用于优化提示、程序和机器学习(ML)工作流的系统通常依赖显式外循环控制器,如进化搜索、多臂老虎机或文本梯度方法。本文提出一个根本不同的问题:这种搜索策略能在多大程度上被单个使用工具的智能体内化?我们提出ReASearch,一个推理驱动优化的统一框架,其中智能体自主决定要评估的内容、如何诊断失败、要进行哪些编辑,以及何时验证或重启。智能体并非仅作为受手工设计启发式引导的提案生成器,而是通过持久记忆主动分析结果、分配预算并在长期过程中优化策略。借助共享的智能体循环和特定领域工具,ReASearch采用完全相同的框架来优化提示、程序和ML工作流。在14项不同任务中,它与专用优化系统具有竞争力且多数情况下表现更优,相比强大的特定领域基线实现了2%至40%的提升,部分情况下还发现了优于此前人类已知最佳结果的解决方案。关键在于,我们观察到通常由显式控制器实现的复杂搜索行为,会自然地从智能体的推理过程中涌现出来。

英文摘要

Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Snowflake(思诺飞克公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑