发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有智能体评估仅看结果的缺口,提出双评估基准ClawTrack,含320个任务及过程评分器,评估21个模型后发现其可归因推理、验证结果验证瓶颈等,助力智能体改进。
AI 中文摘要
当基于大语言模型(LLM)的智能体部署在复杂的多步骤工作流中时,出现了一个关键的评估缺口:大多数现有基准仅判断最终结果,无法区分可靠的推理与侥幸的成功,也无法将失败归因于特定的过程缺陷,这阻碍了长程任务中的归因。本研究提出了ClawTrack,这是一个双评估基准,可同时测量智能体达成的结果(任务得分)与达成的过程(过程得分)。ClawTrack包含8个领域的320个任务,搭配25个以上的确定性模拟服务。过程评分器(Process Grader)沿四个维度对每个推理轮次进行评分,四个维度为目标一致性、效率、信息利用和结果验证,由12541个任务特定的评分标准项作为锚点。对21个模型开展超过16000次试验的评估发现:(1)过程得分可有效将成功与失败归因于特定推理维度,过滤掉仅靠结果评估无法察觉的侥幸通过;(2)四个维度具有互补性,其中结果验证是系统性瓶颈;(3)该框架对不同评判大语言模型的选择具有鲁棒性;(4)基于过程的轨迹过滤可在不同模型规模上实现一致的训练后改进。
英文摘要
As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks. In this work, we present ClawTrack, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task Score) and how it achieves it (Process Score). ClawTrack comprises 320 tasks across 8 domains with 25+ deterministic mock services. A Process Grader scores each reasoning turn along four dimensions (goal alignment, efficiency, information utilization, and result verification), anchored by 12,541 task-specific rubric items. Evaluating 21 models over 16,000+ trials, we find that: (1) process scores effectively attribute success and failure to specific reasoning dimensions, filtering lucky passes invisible to outcome-only evaluation; (2) the four dimensions are complementary, with result verification as the systematic bottleneck; (3) the framework is robust to evaluator choice across different judge LLMs; and (4) process-based trajectory filtering yields consistent post-training improvements across model scales.