harness优化价值存在于何处?自进化大语言模型智能体中的局部增益与预算拆分陷阱
Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究提出HARNESSEVO将harness分解为四个槽,发现harness优化价值集中在反思/控制槽,均匀拆分预算有害,集中预算于高贡献控制槽可提升ALFWorld任务性能,WebShop任务无此现象。
中文摘要 AI 辅助
越来越多的研究通过进化大型语言模型(LLM)的harness(围绕模型的文本支架,包括角色、策略、格式规则和控制启发式),将冻结的LLM改进为智能体。现有的反思式提示进化方法通常将该harness优化为单个扁平字符串。我们转而探究优化价值实际存在的位置。我们引入HARNESSEVO,将harness分解为四个可单独进化的槽:角色、任务策略、工具/格式规则以及反思/控制。在等预算设置下使用相同的反思式优化器,我们将该分解与留一加入和留一剔除归因法配对,以测量每个槽的贡献。在使用冻结的7B主干模型的ALFWorld任务上,HARNESSEVO的整体二元成功率并未显著优于原始harness或扁平字符串进化:分别为0.657,而原始和扁平进化均为0.642。然而,槽级分析显示,几乎所有有用的优化价值都集中在反思/控制槽,其留一加入增益为+0.119,其他槽单独无增益。我们进一步表明,均匀预算拆分是有害的:在四个槽中分配64次rollout,每个槽仅得16次,低于优化器的有效搜索下限,导致每个槽都冻结在其空种子状态。将预算集中在高贡献的控制槽可恢复损失的增益,在仅为拆分预算一半的情况下达到0.761。该效果具有任务依赖性:在WebShop任务上,所有槽都冻结为空,所有方法表现相当,表明确实不存在反复出现的、可语言化的控制失败,而非预算不足。总体而言,我们的结果表明harness价值是局部化的,均匀预算拆分可能具有主动危害,且在结构化智能体进化前应先进行信用分配。
英文摘要
A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflective prompt-evolution methods usually optimize this harness as one flat string. We instead ask where the optimization value actually resides. We introduce HARNESSEVO, which decomposes the harness into four separately evolvable slots: role, task-strategy, tool/format-rules, and reflection/control. Using the same reflective optimizer under an iso-budget setting, we pair this decomposition with leave-one-in and leave-one-out attribution to measure the contribution of each slot. On ALFWorld with a frozen 7B backbone, HARNESSEVO does not significantly improve the overall binary success rate over either the stock harness or flat-string evolution: 0.657 versus 0.642 and 0.642, respectively. However, the slot-level analysis reveals that nearly all useful optimization value is localized in the reflection/control slot, which achieves a leave-one-in gain of +0.119. The other slots are individually null. We further show that uniform budget splitting is harmful: allocating 64 rollouts across four slots leaves only 16 per slot, below the optimizer's effective search floor, causing every slot to freeze at its empty seed. Concentrating the budget on the high-credit control slot recovers the lost gain, reaching 0.761 with half the split budget. The effect is task-contingent. On WebShop, all slots freeze empty and all methods tie, indicating a genuine absence of recurrent, verbalizable control failures rather than budget starvation. Overall, our results suggest that harness value is localized, uniform budget splitting can be actively harmful, and credit assignment should precede structured agent-evolution.
发表机构
- Universiti Malaya(马来亚大学)
- Universiti Sains Malaysia(马来西亚理科大学)
- Universiti Putra Malaysia(马来西亚博特拉大学)
- Monash University Malaysia(莫纳什大学马来西亚分校)
机构由 AI 辅助整理,请以论文原文为准。