arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38143cs.AIcs.CLcs.LG

学习测试时AI4AI智能体框架设计中的元技能

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang, Heng Ji

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出元技能方法,使Builder从执行反馈中学习构建智能体环境的原则,在基准测试上显著提升性能,并指向系统级自我改进路径。

中文摘要 AI 辅助

智能体性能既取决于推理能力,也取决于其运行环境。我们研究测试时AI-for-AI问题,探讨Builder如何在两个模型权重均保持不变的情况下,学习为Target构建更好的执行环境。为使Builder的经验可复用,我们引入元技能(Meta-Skill):即规定何时需要支持以及提供何种资源的原则。Builder从Target在开发集上的执行反馈中学习这些原则,然后利用冻结的技能库为未见任务构建框架。在Harness-Bench和NewtonBench上,全库元技能相比无技能构建,宏平均性能提升了8.95个百分点;相比将同一技能库直接提供给Target,提升了12.02个百分点。这些结果凸显了将经验转化为可执行支持的价值。当同一模型同时担任两个角色时获得的增益,进一步指明了通过学会构建更好的环境实现系统级自我改进的路径。

英文摘要

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

发表机构

  • Apodex
  • University of Illinois Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑