arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25730cs.CR

从裁决到诊断:对拉取请求的可归因安全审查

From Verdict to Diagnosis: Attributable Security Review of Pull Requests

Zhuo Chen, Boyang Wang, Xiyue Zhang, Xiaoyun Xu, Ahmad-Reza Sadeghi, Stjepan Picek, Lichao Wu

中文总结 AI 辅助

研究针对自动化PR审查的裁决-诊断差距,构建MalPR-Bench基准,提出PRGuard审查工具,发现仅基于裁决的评估会高估自动化审查的安全价值。

中文摘要 AI 辅助

自动化代码审查工具越来越多地被用作拉取请求(PR)的准入关卡,但现有评估仅衡量它们是否会阻止恶意变更。阻止可能由不相关问题触发,而非使PR不安全的漏洞;修复报告的问题可能会让目标缺陷仍可被利用,我们将这种差异称为裁决-诊断(VD)差距。我们提出MalPR-Bench,这是一个基于机制的基准,包含44个仓库、8种语言族的89个恶意PR和50个配对的良性对照。每个恶意案例都有预先提交的规则,指定目标漏洞、可接受的机制描述、所需的仓库证据,以及不被认可的偏离目标的发现。审查分别按裁决正确性、目标漏洞识别和证据验证评分;可归因的阻止需要满足所有三个条件。我们引入PRGuard,这是一种可归因的PR安全审查工具,它构建候选漏洞并使用确定性、非执行工具和有限检索来验证其针对仓库证据的前提。在31个常见覆盖的保留恶意PR中,PRGuard和CodeRabbit的阻止总数相似(22/31对24/31),但PRGuard识别出22个目标漏洞,而CodeRabbit为16个,差异为1.38倍。在14个缺失类型的案例中,两者都阻止了9个,而PRGuard识别出9个目标,CodeRabbit为3个。当所需证据位于已修改文件内时,CodeRabbit识别出16/24个目标,当验证需要文件外的证据时则为0/7。最后,PRGuard在5个项目中发现了12个以前未披露的、有概念验证支持的漏洞。PRGuard/DeepSeek和CodeRabbit都阻止了10/12个发现PR,但分别产生了10/12和4/12个可归因阻止。因此,仅基于裁决的评估可能会大大高估自动化审查的安全价值。

英文摘要

Automated code reviewers are increasingly used as gates on pull requests (PRs), yet evaluations measure whether they block a malicious change. A block may be triggered by an unrelated issue rather than the vulnerability that makes the PR unsafe; fixing the reported issue can leave the target defect exploitable. We call this discrepancy the Verdict-Diagnosis (VD) gap. We present MalPR-Bench, a mechanism-grounded benchmark of 89 malicious PRs and 50 paired benign controls across 44 repositories and eight language families. Each malicious case has a pre-committed rubric specifying the target vulnerability, accepted mechanism descriptions, required repository evidence, and off-target findings receiving no credit. Reviews are scored separately for verdict correctness, target-vulnerability identification, and evidence validation; an attributable block requires all three. We introduce PRGuard, an attributable PR security reviewer that constructs candidate vulnerabilities and validates their premises against repository evidence using deterministic, non-executing tools and bounded retrieval. Across 31 common-coverage held-out malicious PRs, PRGuard and CodeRabbit produce similar blocking totals (22/31 vs. 24/31), but PRGuard identifies 22 target vulnerabilities versus 16 for CodeRabbit, a 1.38x difference. On 14 absence-type cases, both block 9, while PRGuard identifies 9 targets versus 3. CodeRabbit identifies 16/24 targets when required evidence lies within touched files and 0/7 when validation requires evidence outside them. Finally, PRGuard uncovers twelve previously undisclosed, proof-of-concept-backed vulnerabilities across five projects. PRGuard/DeepSeek and CodeRabbit both block 10/12 discovery PRs, but produce 10/12 and 4/12 attributable blocks, respectively. Thus, verdict-only evaluation can substantially overstate the security value of automated review.

↑