ParanoiaEval:基准测试智能体编码中的不必要防御性工作
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
浏览论文内容
中文总结 AI 辅助
ParanoiaEval是首个统一评估编码智能体风险处理能力的基准,基于回避-转移-缓解-接受框架,含200个证据控制任务对,发现不必要的防御性工作占11.2%-58.7%,且风险处理是独立于任务能力的关键维度。
中文摘要 AI 辅助
随着编码智能体日益自主地承担现实世界的工作,判断其风险处理是否合理已变得至关重要。现有工作从不同视角评估相关智能体行为,但缺乏一个系统化的框架来统一这些行为。为填补这一空白,我们引入了ParanoiaEval,这是首个用于统一评估编码智能体风险处理能力的基准。基于软件工程风险管理中成熟的回避-转移-缓解-接受框架,ParanoiaEval将其4种基本处理方式操作化应用于编码智能体场景,并包含200个证据控制的仓库级任务对,每对任务仅在定义处理的证据上有所不同。我们进一步引入了针对风险处理违规和证据响应性的专用指标,并使用经过人工校准的智能体评判器进行可靠评估。对8个代表性模型的大规模实验和一项事后人工研究揭示:(I)尽管有明确证据,不必要的风险处理仍发生在11.2%-58.7%的运行中,且在不同智能体配置间存在显著差异;(II)更强的任务能力并不确保更适当的风险处理,而处理违规会显著损害开发者体验,从而确立风险处理为一个独立的能力维度;(III)智能体表现出与既定风险管理发现一致的系统性模式,表明来自人类实践的知识可以指导这一能力的诊断和改进。
英文摘要
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
发表机构
- New York University(纽约大学)
- New York University Abu Dhabi(纽约大学阿布扎比分校)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。