arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EarlyEval:通过早期结果预测实现更经济的智能体评估

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Yuling Shi, Zhensu Sun, Junsen Dong, Chengcheng Wan, David Lo, Xiaodong Gu

arXiv 2609.02783首次发表:更新:

发表机构

Shanghai Jiao Tong University; Singapore Management University; East China Normal University; Shanghai Innovation Institute(上海交通大学; 新加坡管理大学; 华东师范大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出EarlyEval框架,通过训练LightGBM分类器预测智能体结果,提前终止执行以降低LLM智能体评估成本,在多个基准上实现了显著的成本削减且解决率波动极小。

AI 中文摘要

评估大型语言模型(LLM)智能体对指导其开发至关重要,但评估成本已高得令人望而却步:前沿模型对一个智能体基准进行一次评估就需花费数百至数千美元,且在迭代开发周期中需重复支付该成本。以往研究以基准蒸馏为核心,减少评估任务数量但未降低每个保留任务的执行成本。本研究引入早期结果预测,这是一种互补的效率优化方向,可在每个任务内降低成本。核心洞见为:智能体的最终结果通常在执行完成前很久就可从其中间行为中明确。我们将该思路实例化为EarlyEval,这是一个轻量级框架,它基于行为、文本及参考解决方案特征训练一对LightGBM成功与失败分类器,并在任一分类器达到校准置信阈值时立即停止智能体运行,仅添加可忽略的每步开销。在三个基准(SWE-bench Verified、TerminalBench和Toolathlon)上,EarlyEval可消除13%-26%的智能体步骤,最多减少44.1%的输入令牌和29.4%的输出令牌,预测准确率达89%-97%,同时使每个智能体的解决率平均仅波动1至2个百分点。

英文摘要

Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome is often evident from its intermediate behavior well before execution completes. We instantiate this idea in EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features, and halts an agent run the moment either classifier crosses a calibrated confidence threshold, adding negligible per-step overhead. Across three benchmarks, SWE-bench Verified, TerminalBench, and Toolathlon, EarlyEval can eliminate 13%-26% of agent steps and up to 44.1% input tokens and 29.4% output tokens at 89%-97% prediction accuracy, while perturbing per-agent resolve rates by only one to two percentage points on average.

CommentsCode and data available at https://github.com/inphotoo/earlyeval

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑