arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38762cs.SEcs.AI

Adaptive-GEPA:让您的工具适配异构请求

Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

Tianyu Chen, Yasi Zhang, Ruiyi Wang, Xinran Zhao, Taoran Li, Mingyuan Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

Adaptive-GEPA通过联合演化路由器与专家程序库,在单一搜索预算下自动划分并解决异构请求,显著提升测试分数。

中文摘要 AI 辅助

反射式优化器(如GEPA)通过执行轨迹和评估器反馈改进语言模型提示;全程序扩展还可以重写工具和控制流。在实践中,用户向同一端点提交异构请求,其有效解决方案需要不同的工具、推理模式和控制流。优化一个共享程序会使这种工作划分在源代码搜索中隐式化,而按请求族分别优化程序则事先固定了划分。我们提出Adaptive-GEPA,它同时学习如何划分请求以及如何解决它们。它在一个搜索预算下演化出一个路由器和一个专家程序库。路由器的指令、每个专家的描述及其程序代码都是纯文本、人类可读的,并根据反馈进行编辑。为了合并分支,它根据专家处理的请求对其进行对齐,并连同程序一起继承描述。在四个任务族的固定混合上,报告的Qwen3-8B运行在未向路由器或反射模型提供族标签的情况下演化出四个专家;其路由在所有651个测试请求上与任务划分匹配。其族平均测试分数(×100)从52.6提升到70.6,而GEPA的全程序适配器为62.5,GRPO在名义预算18,000次评分调用下为54.0。这些计数并不等同于总计算量。图1总结了学习曲线、最终测试分数和路由一致性。

英文摘要

Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and control flow. Optimizing one shared program leaves this division of work implicit in source-code search, while optimizing a separate program per request family fixes it beforehand. We introduce Adaptive-GEPA, which learns both how to divide requests and how to solve them. It evolves a router and a library of specialist programs under one search budget. The router's instructions, each specialist's description, and its program code are plain, human-readable text, edited from feedback. To combine branches, it aligns specialists by the requests they handle and inherits descriptions together with programs. On a fixed mixture of four task families, the reported Qwen3-8B run evolves four experts without supplying family labels to the router or reflection model; its routing matches the task partition on all 651 test requests. Its family-mean test score (x100) rises from 52.6 to 70.6, compared with 62.5 for GEPA's full-program adapter and 54.0 for GRPO at a nominal budget of 18,000 scored calls. These counts do not equate total compute. Figure 1 summarizes the learning curves, final test scores, and routing agreement.

发表机构

  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Shenzhen Institutes of Advanced Technology(深圳先进技术研究院)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑