发表机构
Explyt; St. Petersburg Department of the Steklov Institute of Mathematics; St. Petersburg State University(Explyt; 圣彼得堡斯捷克洛夫数学研究所圣彼得堡分部; 圣彼得堡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍用于编码智能体评估的AgentLens基准,结合形式验证、大语言模型编写的轨迹审查和并排比较来评估智能体整个轨迹,不仅能排名,还可用于诊断模型行为、比较智能体版本及发现产品回归问题,且已开源。
AI 中文摘要
我们提出了AgentLens,这是一个用于交互式代码智能体的生产评估基准。大多数代码智能体基准测试将一次运行简化为单个结果——任务是否通过?但实际使用这些智能体的人会经历整个轨迹:智能体如何遵循指令、使用工具、验证自身工作、从错误中恢复以及在此过程中与他们交流。AgentLens评估整个轨迹。它将存在客观检查的形式验证与由大语言模型编写的轨迹审查和并排比较相结合,以便每次运行都能产生关于分数为何如此的可读解释。这使得AgentLens不仅可用于对模型进行排名:我们还用它来诊断模型行为、比较我们自己智能体的连续版本,并在夜间评估管道中发现产品回归问题。我们将该基准作为开源项目发布在这个https网址。
英文摘要
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.