发表机构
University of Tübingen; ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center; École Polytechnique; Institut Polytechnique de Paris; CNRS; KAIST AI(蒂宾根大学; ELLIS研究所蒂宾根; 马克斯·普朗克智能系统研究所; 蒂宾根人工智能中心; 巴黎综合理工学院; 巴黎综合理工学院; 法国国家科学研究中心; 韩国科学技术院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出BOTTLED基准,检验LLM智能体能否将通用能力转化为廉价可扩展的特定任务解决方案,发现零样本性能不等于装瓶能力,但装瓶可大幅降低成本并保持较高性能。
AI 中文摘要
大型语言模型(LLM)能够解决许多狭窄的任务,但针对数百万个相关实例分别进行查询可能代价高昂。LLM智能体能否自主地为这类工作负载创建更廉价的解决方案?我们将这种能力称为“装瓶”(bottling):即将通用能力转化为在答案质量与摊销成本之间取得平衡的特定任务解决方案的能力。我们引入了BOTTLED基准,在该基准中,智能体接收整个未标注的工作负载,并必须在固定的时间、计算和LLM API预算内完成。智能体自行选择方法,例如训练一个小型模型或编写一个可复用的程序。在十个模型和三个任务上,我们发现,强大的零样本任务性能并不能可靠地转化为强大的装瓶能力。零样本得分相似的模型在装瓶后可能表现出显著差异,且60次装瓶运行中有48次的得分低于其模型零样本性能95%置信区间的下限。此外,60次运行中有31次的表现不及使用相同令牌预算的两个小型模型蒸馏基线中较强的一个。尽管如此,装瓶可以带来可观的节省:在查询-产品相关性分类任务上,Opus 5保留了其零样本宏F1约82%的性能,而报告成本约为原来的657分之一。装瓶还与Jev(一种专为廉价、重复推理而构建的“系统一”模型)具有竞争力:在同一任务上,Opus 5恢复了Jev宏F1约94%的性能,而预计的全工作负载成本仅为Jev的四分之一。BOTTLED为评估和改进智能体将有限资源投资于可复用解决方案以处理大型重复工作负载的能力提供了基础。
英文摘要
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.