arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36730cs.AIcs.CLcs.SE

智能体能否为智能体设计库?

Can Agents Design Libraries for Agents?

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Snorkel AI
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt

AI总结:

提出LibraryDesignBench两阶段基准,衡量智能体为其他智能体设计库的能力,通过下游智能体程序的正确性和简洁性评估,发现智能体编写的库僵化或难用导致下游重实现,指导性指南和子智能体测试可改善下游得分和程序简洁性。

AI中文摘要:

智能体越来越依赖于其他智能体编写的代码,并且它们倾向于重新实现而非复用,这导致后续智能体必须处理的代码库不断膨胀。为了衡量智能体为其他智能体设计库的能力,我们引入了LibraryDesignBench,这是一个两阶段基准测试,其中智能体根据规范实现一个功能完备的库,该规范定义了所需能力和潜在用例,但不规定具体设计。我们通过三个来自不同模型家族的用户智能体所编写程序的正确性和简洁性来评估该库。该基准涵盖四种语言中15个库设计任务中的242个专家验证的编程问题。在15个任务中的11个上,智能体设计者复现了人工编写的生产库的抽象。下游智能体同样采用智能体编写和人工编写的库,但未能充分利用它们,重新实现了库已提供的能力。我们的失败分析发现,下游智能体编写额外代码的主要原因是智能体编写的库过于僵化或难以使用,而非能力缺失。我们还尝试为设计者提供更具指导性的、以智能体为先的指南,并让它们使用子智能体测试其库;这改善了下游得分并产生了更简洁的程序。LibraryDesignBench既为评估面向智能体用户的库设计实践提供了测试平台,也提供了一个改善下游复用的初始设计基线。

英文摘要:

Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.

补充信息

↑