arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WHALE:一种用于联合 harness-权重优化的简单方案

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee

arXiv 2609.00196首次发表:更新:

发表机构

KRAFTON; KAIST; Stanford University(克拉夫顿公司; 韩国科学技术院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体性能受模型权重与 harness 代码单独优化瓶颈的问题,提出 WHALE 交替优化方案,结合在线拒绝采样微调与 Meta-Harness,在多领域 Qwen 智能体上较基线方法提升了准确率与 rollout 效率。

AI 中文摘要

智能体的性能同时取决于模型参数和管理上下文与控制流的可执行 harness 代码。单独优化任一组件都会导致系统被另一固定组件限制:权重更新会改变有效的 harness,而 harness 更新会改变模型能力的可展示性。现有的联合适配方法会优化权重和文本提示,但会固定更广泛的 harness。我们提出 Weight-Harness Alternating LEarning(WHALE,权重-harness 交替学习),一种交替两个阶段的简单方案:在当前 harness 下更新模型,再在更新后的模型下搜索更好的 harness。我们分别用在线拒绝采样微调与 Meta-Harness 实现这两个阶段。何时切换是关键设计选择:为了分离真实改进与噪声,且不会对变化的对应组件过度优化,WHALE 使用固定阶段时长或基于训练信号的自适应耐心规则。在三个领域(搜索式问答、数学推理、国际象棋谜题)中使用 Qwen3.5-2B/4B 智能体时,WHALE 的最佳 mean@8 准确率比仅权重优化、仅 harness 优化及 Fast-Slow Training 高出 4.15-24.38 个百分点。任一组件都可能成为瓶颈:在 SearchQA 中,harness 搜索用少得多的 rollout 匹配仅权重优化的峰值准确率,但仅在权重更新后才提升数学准确率;小的交替更新在准确率和 rollout 成本上也优于分阶段的“先权重后 harness”优化。代码可在该 https URL 获取。

英文摘要

Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. Existing joint-adaptation methods optimize weights and textual prompts but leave the broader harness fixed. We propose Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model. We instantiate these two phases with online rejection-sampling fine-tuning and Meta-Harness, respectively. When to switch is a key design choice: to separate real improvements from noise without over-optimizing against a changing counterpart, WHALE uses either fixed phase durations or an adaptive patience rule over training signals. Using Qwen3.5-2B/4B agents across three domains (search question answering, mathematical reasoning, and chess puzzles), WHALE outperforms weight-only, harness-only, and Fast-Slow Training by 4.15-24.38 percentage points in best mean@8 accuracy. Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update. Small interleaved updates also outperform stagewise weight-then-harness optimization in accuracy and rollout cost. The code is available at https://github.com/krafton-ai/WHALE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑