RobustSGPO:智能体工具链演化的搜索空间控制
RobustSGPO: Search-Space Control for Agent Harness Evolution
浏览论文内容
中文总结 AI 辅助
RobustSGPO通过明确编辑范围与操作、保留快照和搜索空间控制,在AgentX工作流中显著提升任务完成率和测试质量,优于固定权限调度。
中文摘要 AI 辅助
基于语义梯度的提示优化(SGPO)利用执行反馈改进智能体工具链,但其局部更新规则未解决编辑范围与操作的选择问题。我们提出RobustSGPO,它明确指定所需编辑,构建并检查补丁,并从当前方案或保留的快照中继续搜索。我们在AgentX头脑风暴工作流中,使用120个任务、95次运行和7,350次候选尝试,评估了权限调度、累积控制和任务族迁移。周期性$1\ o2\ o3$调度比固定最大权限高出0.28个测试分数点。在2000万token预算下,RobustSGPO将30个保留任务的完成率从60.0%提升至80.0%,并将测试质量从3.77提升至4.14。类别保留减少了迁移后源任务的性能下降,而随机保留则达到更高的目标端点。搜索空间控制通过可执行编辑和替代起点提升质量,并带来可测量的保留开销。
英文摘要
Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks the patch, and continues search from either the incumbent or retained snapshots. We evaluate permission scheduling, cumulative controls, and task-family transfer in the AgentX brainstorming workflow using 120 tasks, 95 runs, and 7,350 candidate attempts. Periodic $1\to2\to3$ scheduling exceeds fixed maximum permission by 0.28 test-score points. RobustSGPO increases completion on 30 held-out tasks from 60.0% to 80.0% and improves test quality from 3.77 to 4.14 under a 20-million-token budget. Category retention reduces source-task degradation after a shift, whereas random retention reaches a higher destination endpoint. Search-space control benefits quality through executable edits and alternative starting points, with measurable retention overhead.
发表机构
- Wuhan University(武汉大学)
- Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。