AI 中文总结
该研究提出ParEvalLayer决策层,可基于部分任务结果判断LLM智能体比较结论,部分基准仅需15%-25%任务结果即可得出与完整评估一致的决策。
AI 中文摘要
大语言模型智能体评估通常在完整基准测试运行完成前很久就会产生任务结果。部分得分很诱人,但它无法表明观察到的任务是否与完整评估得出相同结论。早期任务可能会省略基准测试的重要部分,先运行成本较低的任务可能会扭曲观察到的样本,仅决定简单对的规则可能看似准确,却留下许多未解决的比较。我们提出ParEvalLayer,这是一个决策层,它读取两个智能体系统的配对结果和预先选定的比较策略。对于每次部分运行,它记录被测试智能体系统是否达到所需的更优程度、未达到该程度、需要更多证据,还是应弃权(不执行)。我们通过重放已完成的公开基准数据来评估ParEvalLayer,就好像每次评估都提前停止了一样。在每一步,ParEvalLayer仅使用迄今为止观察到的结果应用该策略;如果它得出两种比较判断之一,我们会检查该判断是否与同一系统对的完整数据匹配。使用主要比较规则,三个公开基准在仅观察到15%至25%的任务结果后,就得出了与完整评估相同的决策。其他基准需要更多任务结果。这种差异表明为何仅部分得分不够:报告还应说明决策规则以及仍未得出决策的比较数量。
英文摘要
LLM-agent evaluations often produce task outcomes long before the full benchmark run is complete. A partial score is tempting to report, but it does not show whether the observed tasks support the same conclusion as the completed evaluation. Early tasks can omit important parts of a benchmark, running cheaper tasks first can distort the observed sample, and a rule that decides only easy pairs can appear accurate while leaving many comparisons unresolved. We introduce ParEvalLayer, a decision layer that reads paired outcomes for two agent systems and a comparison policy chosen in advance. For each partial run, it records whether the tested agent system is better by the required amount, is not better by that amount, needs more evidence, or should abstain. We evaluate ParEvalLayer by replaying completed public benchmark data as if each evaluation had stopped earlier. At each point, ParEvalLayer applies the policy using only the outcomes observed so far; if it reaches one of the two comparison judgments, we check whether that judgment matches the completed data for the same system pair. With the main comparison rule, three of the public benchmarks reach the same decision as the completed evaluation after observing only 15% to 25% of task outcomes. Other benchmarks require more task outcomes. This variation shows why a partial score alone is not enough: reports should also state the decision rule and how many comparisons remain without a decision.
CommentsAccepted at the 2026 ACM International Conference on AI-ML Systems (AIMLSystems)