RealSWE:面向真实用户请求的编码智能体组合式评估
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
- Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究针对编码智能体评估基准与真实用户请求的差异,构建了新基准sys并评估7种LLM,发现真实输入降低解决率且改变模型排名,明确陈述期望行为和动机可提升LLM软件工程性能。
中文摘要 AI 辅助
当前编码智能体通常在SWE-bench系列基准上进行评估,该基准的任务由精心整理的GitHub issues构建,具有冗长、结构化且信息丰富的特点。然而,真实用户请求通常更简短且结构化程度更低。为刻画这一差距,我们定义了包含6个类别的信息分类法和4个语言风格维度,并将其应用于SWE-chat的真实用户提示,以及SWE-bench Verified和Pro的问题陈述。我们发现,仅包含问题陈述(单独或搭配有限额外上下文)的请求占真实提示的88%,但仅占基准问题的7%;此外,87%的真实提示为非正式写作,而94%的基准问题为正式写作。基于这些观察,我们引入sys——源自SWE-bench Verified和Pro的381个多变体任务家族,每个家族内的变体共享相同的底层任务和黄金补丁,仅在信息构成和语言风格上存在差异。使用sys评估7种当代大语言模型(LLM),我们发现:i)真实输入使解决率平均降低6.4个百分点,且会改变模型排名;受控分析进一步显示,ii)包含期望行为和动机显著影响性能,而环境信息和复现步骤仅增加token数量,无明显收益;iii)语言风格仅产生较小的、依赖于模型的影响。这些发现为用户和智能体提供了可操作的指导:明确陈述期望行为和动机(大多数真实提示未包含)可大幅提升LLM的软件工程性能。
英文摘要
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of real prompts are casually written whereas 94% of benchmark problems are formal. Guided by these observations, we introduce RealSWE, 381 multi-variant task families derived from SWE-bench Verified and Pro. Variants within each family share the same underlying task and gold patch while differing only in information composition and linguistic style. Evaluating seven contemporary LLMs with RealSWE, we find that i) realistic inputs reduce resolution rates by 6.4 pp on average and can change model rankings. Controlled analysis further shows that ii) including Desired Behavior and Motivation significantly affects performance, whereas Environment Information and Reproduction Steps merely add tokens without measurable benefit; iii) linguistic style has only small, model-dependent effects. These findings provide actionable guidance for users and agents: explicitly stating the desired behavior and motivation, which most real prompts omit, substantially improves the LLM's software engineering performance.