arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Repo-To-Skill:将GitHub仓库提炼为AI4AI技能

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu

arXiv 2609.02749首次发表:更新:

发表机构

Beijing Academy of Artificial Intelligence; University of Science and Technology of China; Renmin University of China; Hong Kong Polytechnic University(北京人工智能研究院; 中国科学技术大学; 中国人民大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DisCo智能体,可将GitHub仓库提炼为技能,构建含5000+技能的AREX-Skill Library,使研究智能体在多个基准测试中得分显著提升。

AI 中文摘要

自主智能体正开始端到端地开展机器学习(ML)研究。这些智能体将模型主干与规划、执行、记忆及验证工具相结合,但该架构仍将特定领域的专门知识排除在智能体之外。我们将这一缺失层称为操作知识,即区分知晓某一方法与使其可行的专门知识。这类知识并未在该领域消失,它以面向人类读者的形式出现在仓库和论文中,但规模过大,无法在任务执行期间加载。一旦提炼为紧凑且经过验证的技能,这类知识便可在不同任务间复用,而非在每次运行时重新发现。我们提出DisCo,一种由技能驱动的研究智能体,它可在研究过程中创建并使用技能。其提炼过程以两种互补形式开展:与任务无关的形式,将该领域广泛使用的仓库浓缩为可复用技能;面向任务的形式,生成具体任务所需的技能。前者应用于开放生态系统,构建出AREX-Skill Library,其中包含从1000个广泛使用的ML仓库中提炼的5000多个经过验证的技能,这些技能被组织为20个领域和178个能力族。在固定GPT-5.5主干、研究工具及下游执行预算的情况下,配备技能的研究智能体在MLE-bench上的得分比未配备技能的同一智能体高134.3%,在PaperBench上高34.4%,在FrontierCS上高9.2%,在PassNet上高14.0%。这些增益来自在固定设置下添加了提炼的操作上下文。

英文摘要

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.

Comments48 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑