发表机构
Haas School of Business(哈斯商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CentaurBench基准测试显示LLM的自动化能力与辅助能力仅适度相关,多数自动化优胜者在增强任务中表现不佳,辅助效果并非始终为正,表明需按模型在多智能体系统中的角色评估模型。
AI 中文摘要
大多数大型语言模型(LLM)基准测试根据模型自动化工作任务的能力对其进行排名。然而在实践中,模型常被用于辅助其他(人类或LLM)智能体。因此,驱动模型选择的问题不仅是哪个模型能产出最佳结果,更是哪个模型能最大程度提升另一个(能力较弱的)智能体的工作表现。我们提出了一个统一框架,用于评估模型自动化和增强其他智能体表现的能力。在7项基于经济场景的现实任务中,辅助模型为标准化的低能力工作模型编写辅助文本,由后者生成交付成果;在自动化模式下,辅助模型直接生成结果。通过带有任务特定规则的LLM评审小组进行盲法成对比较来对结果打分,重复10次实验。两种模式下的排名仅呈适度相关,且自动化模式的优胜者在7项任务中有5项在增强模式下表现更差。辅助效果并非始终为正:在3项任务中,未获辅助的工作模型排名优于所有辅助条件,且仅1个模型的指导平均而言优于无指导。这些结果表明,自动化能力是辅助质量的不完整代理,为根据模型在人机及多智能体系统中扮演的角色评估模型的基准测试提供了动机。
英文摘要
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Comments46 pages, 15 figures