EcoAgent-Bench:评估预算约束下大语言模型智能体的经济决策能力
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
浏览论文内容
中文总结 AI 辅助
本文提出EcoAgent-Bench基准,评估大语言模型智能体在预算约束下的经济决策能力,发现现有智能体在该任务上表现不佳,发布了相关研究资源。
中文摘要 AI 辅助
智能体基准通常仅衡量任务完成情况,将资源使用视为辅助统计指标。但在实际部署中,在本地查询、广泛搜索、复合研究工具、更强模型或人工介入之间的选择本身就是任务的一部分。本文提出EcoAgent-Bench,其中每个任务都指定了带价格的动作和明确的预算。该基准包含304个源自真实场景的任务,涵盖从GAIA、HotpotQA和MuSiQue改编的5个任务类别,测试四种决策:避免不必要的人工介入、本地证据不足时进行人工介入、选择模型层级以及在前提不成立时弃权(不执行)。我们在工具-API和工作区-CLI设置下评估了7个大语言模型智能体,同时评估了4个基于预言机的脚本控制模型。微平均准确率倾向于单边策略:始终选择人工介入的控制模型实现了较高的微平均成功率,但在以节约为导向的任务中失败。因此我们还报告了经济一致性得分(升级导向和节约导向任务组准确率的较小值),该得分揭示了这种失败。工具-API智能体仅达到3.9%-24.0%的微严格成功率(经济一致性最高为7.3%),往往要么在应介入前停止,要么在低成本任务上超支。阈值交叉预算扫描将GPT-5.4的介入率从0%仅提升至3%。这些结果表明,预算下的完成情况和经济动作选择是两种不同的属性。我们发布了任务包、转换管道、冻结评估环境以及研究这两种属性所需的完整性绑定结果制品。
英文摘要
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4's escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.