AI 中文总结
研究在模型-工具协同进化中,通过递归工具自我改进(RHI)优化用户构建工具,以提高执行轨迹质量和智能体性能。RHI将工具表示为提示级规范,经成对反馈迭代改进,在多任务中提升了低推理能力智能体性能,降低推理成本。
AI 中文摘要
在模型-工具协同进化中,工具不仅是推理时的支架,还是数据生成组件,其执行轨迹可塑造未来基础模型。这激发了工具在循环学习,即优化工具以提高即时智能体性能和用于未来模型训练的轨迹质量。但持续更新供应商构建的支架成本高且劳动密集。因此研究以任务特定方式优化用户构建的工具能否提高执行轨迹质量,同时保持计算轻量级且只需几次更新迭代。为此引入递归工具自我改进(RHI),将工具表示为智能体循环的提示级规范,并通过对自身修订历史的成对反馈迭代改进。在30个合成机器学习研究任务中,几次RHI迭代足以大幅提高低推理能力智能体的性能上限,超过相应的最大推理能力设置,同时将推理成本降低多达60%。这些收益主要源于通过更有效的智能体间信息流改进特定任务上下文管理,而非更长的推理轨迹。最后将此行为形式化为RHI隐式优化目标的信息论假设,表明RHI是模型-工具协同进化范式中持续学习的实用算法。
英文摘要
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.
CommentsThis work addresses the first half of the model-harness coevolution loop