AI 中文总结
本研究提出进化元智能体RoboPhD,通过对含9个LLM端点的菜单进行智能体进化,在DS-1000和PaperFindingBench两个任务上,除一个位置外占据所有帕累托前沿槽位,实现各价格点的竞争优势。
AI 中文摘要
考虑一家公司,其针对特定智能体任务调研竞争对手,力求在每个竞争对手的价格点上提供更优的准确率。若一家公司在帕累托意义上主导了竞争对手,那么没有理性客户会有理由选择其他产品。本文展示了一条实现此类能力的路径,即通过对LLMs菜单进行智能体进化,训练样本池最多仅需100个示例。给定一个包含9个LLM端点的定价菜单、任务、目标及API的简要文档、一个简单的种子智能体,以及操作员为每个问题选定的成本目标(通常设定为现有厂商自身的价格),进化元智能体RoboPhD会逐步进化出完整的智能体程序,逐点攻克两个语义差异较大的任务的公开前沿:DS-1000(经执行检查的代码生成)和PaperFindingBench(由LLM评判的科学文献检索)。我们正式提交的评分结果在两个任务的排行榜上占据了除一个位置外的所有帕累托前沿槽位,包括对得分最高和成本最低的竞争点均实现了帕累托主导。
英文摘要
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy elsewhere. This paper shows a path to this kind of capability by evolving multi-LLM Python agents from training pools of at most 100 examples. Given a priced menu of nine LLM endpoints; brief documentation of the task, objective, and API; a simple seed agent; and an operator-chosen per-problem cost target--usually set at an incumbent's own price--RoboPhD, an evolutionary meta-agent, evolves complete agent programs that attack the public frontiers of two semantically dissimilar tasks point by point: DS-1000 (execution-checked code generation) and PaperFindingBench (LLM-judged scientific document retrieval). On public leaderboards for each task, the evolved agents hold every Pareto-frontier slot but one, including Pareto domination of both the top-scoring and the lowest-cost competing points.
CommentsCode at https://github.com/andborth/RoboPhD