arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层自改进:一种面向任务特定可进化智能体控制框架的方法

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Tailin Zhou

arXiv 2608.08466首次发表:更新:

发表机构

HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HSI框架,使冻结LLM的任务特定控制框架可进化,在中等难度任务上取得性能提升,超出模型能力时无增益,为改进冻结LLM智能体提供可行方向。

AI 中文摘要

现代大语言模型(LLM)智能体通常通过手动修改提示词、工具或工作流进行改进,而围绕模型的可执行支架——即控制框架(harness)——在部署后通常被视为固定不变的产物。本研究探讨一种替代方案,其中控制框架是任务特定且可持续进化的:每个任务族维护其专属的控制框架,该框架通过固定的任务注入接缝在迭代间热交换,并利用环境反馈进行重写。我们提出分层自改进(Hierarchical Self-Improvement, HSI)框架,其中单个冻结的LLM模型M在三个分层范围内运行:执行任务的任务控制框架H、重写H的进化器,以及在冻结的外部锚点下重写进化器策略代码的元进化器。开/关思考设计通过在任务执行期间禁用推理、在自我修改期间启用推理,分离出控制框架进化的贡献。HSI受两个因素限制:一是反馈保真度限制,因为进化需要信息丰富的奖励信号来指导选择;二是主干能力限制,因为控制框架的重新设计无法克服冻结模型的局限性。在BALROG数据集上,以DeepSeek-V4-Flash-Preview作为冻结主干,HSI在中等难度任务上相对于初始控制框架取得持续增益:BabyAI任务原始进度百分比提升39.3,Crafter任务提升33.0,TextWorld任务提升25.0,MiniHack任务提升15.0;同时在BabaIsAI子套件的未见过泛化任务中表现出色,在20%的未见过拆分中,BreakStop的最佳测试得分为0.98,GoTo得分为1.00。在超出主干能力的任务(NLE)上,控制框架进化未提供任何改进。这些结果表明,在明确的经验限制下,任务特定的控制框架进化是改进冻结LLM智能体的可行方向。代码可在该https URL获取。

英文摘要

Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is \emph{task-specific and continuously evolvable}: each task family maintains its own harness, which is hot-swapped across iterations through a fixed task-injection seam and rewritten using environment feedback. We introduce \textbf{Hierarchical Self-Improvement (HSI)}, a framework in which a single frozen LLM $M$ operates across three hierarchical scopes: a task harness $H$ that executes tasks, an evolver that rewrites $H$, and a meta-evolver that rewrites the evolver's strategy code under a frozen outer anchor. A thinking-on/off design isolates the contribution of harness evolution by disabling reasoning during task execution while enabling it during self-modification. HSI is bounded by two factors: a \emph{feedback-fidelity bound}, since evolution requires informative reward signals to guide selection, and a \emph{backbone capability bound}, since harness redesign cannot overcome limitations of the frozen model. On BALROG with DeepSeek-V4-Flash-Preview as the frozen backbone, HSI achieves consistent gains over the initial harness on moderate-difficulty tasks ($+39.3$ on BabyAI, $+33.0$ on Crafter, $+25.0$ on TextWorld, and $+15.0$ on MiniHack, all in raw \% Progress), while obtaining strong held-out generalization on BabaIsAI sub-suites ($0.98$ best-test on BreakStop and $1.00$ on GoTo from a $20\%$ unseen split). On tasks beyond the backbone's capability (NLE), harness evolution provides no improvement. These results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits. Code is available at https://github.com/TailinZhou/hsi.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑