ExplainBench:评估智能体生成的代码解释
ExplainBench: Evaluating Code Explanations from Agents
浏览论文内容
中文总结 AI 辅助
该研究针对智能体代码解释缺乏评估基准的问题,提出ExplainBench基准,实验发现其对智能体的排名与SWE-bench Verified不同,还实现了解释审核智能体,可提升智能体解释的可信度。
中文摘要 AI 辅助
大型语言模型(LLM)智能体在软件工程领域得到了快速应用。随着智能体在实际代码生成中发挥越来越大的作用,它们所做的变更规模也越来越大,涉及数十到数百行代码。这使得人工审核智能体的结果越来越不现实,因此开发者转向依赖解释来理解所实施的变更。尽管如此,目前还没有用于评估智能体生成解释可信度的基准。为了弥合这一差距,我们提出了ExplainBench,这是一个用于自动评估编码智能体生成解释的基准。ExplainBench基于这样一种直觉:有信息量的解释应当能让LLM正确回答问题,从而实现对不同智能体之间解释质量的定量比较。基于这一观察,我们构建了一套问题,用于评估解释是否准确描述了:(1)有缺陷代码的预期行为;(2)应用智能体补丁本身的效果。实验首先表明,解释质量是智能体评估的一个独特维度:ExplainBench对智能体的排名与广泛使用的SWE-bench Verified基准不同。对智能体解释质量的更深入分析显示,解释中存在频繁出现的问题,例如解释常常声称补丁是正确的,而实际上并非如此。基于这一见解,我们实现并评估了一个解释审核智能体,该智能体运行额外的测试来验证和优化解释。这个智能体改进了所有被评估智能体的解释,证明智能体的解释可以自动变得更可信。
英文摘要
Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in the actual generation of code, they are making larger changes, spanning tens to hundreds of lines. This makes manual review of agent results increasingly infeasible, leading developers to turn to explanations to understand enacted changes. Despite this, there are no benchmarks that evaluate the trustworthiness of agent-generated explanations. To bridge this gap, we propose ExplainBench, a benchmark to automatically evaluate explanations from coding agents. ExplainBench is based on the intuition that informative explanations should enable an LLM to correctly answer questions, allowing quantitative comparison of explanation quality between agents. With this observation, we construct a suite of questions that evaluates whether explanations accurately describe (1) the intended behavior of buggy code and (2) the effect of applying the agent patch itself. Experiments first reveal that explanation quality is a distinct axis of agent evaluation: ExplainBench ranks agents differently from the widely-used SWE-bench Verified benchmark. A deeper breakdown of explanation quality in agents shows frequent problems in explanations, such that explanations often claim that a patch is correct when it is not. Based on this insight, we implement and evaluate an explanation audit agent which runs additional tests to validate and refine explanations. This agent improved the explanations of all evaluated agents, demonstrating agent explanations can be automatically made more trustworthy.