PRWeaver:针对基于大语言模型的代码审计工具对抗长周期恶意拉取请求的评估
PRWeaver: Evaluating LLM-Based Code Auditors against Long-Horizon Malicious Pull Requests
浏览论文内容
中文总结 AI 辅助
该研究提出PRWeaver基准,评估基于LLM的代码审计工具对长周期恶意PR的可靠性,发现特定评审策略会降低检测率,为优化代码审计工具提供了依据。
中文摘要 AI 辅助
基于大语言模型(LLM)的代码审计工具正越来越多地被集成到拉取请求(PR)工作流中,但人们对其在应对跨仓库演化分布的对抗性变更时的可靠性仍知之甚少。本文提出PRWeaver,这是一个包含来自10个真实仓库的208次经执行验证的攻击的基准,每次攻击均在4种匹配的评审渲染场景下实例化(总计832种渲染场景)。我们评估了3种PR审计智能体,涉及6种审计模型系统。在所有系统中,仅分解攻击变更会使检测率最多下降5个百分点,这表明仅提交边界无法解释规避行为。相比之下,当N=16时的按PR交织策略和相干载体融合策略分别使检测率降低5-13个百分点和10-18个百分点。在N=24的全窗口评审下,检测率降至16%-22%,而在按PR评审下则为50%-60%。这些结果表明,仅访问仓库历史是不够的:当良性变更与恶意变更共同占据审计工具的活跃评审上下文,或当所述目的看似能合理解释包含攻击的差异时,隐藏行为会最为有效。
英文摘要
LLM-based code auditors are increasingly integrated into pull-request (PR) workflows, yet their reliability against adversarial changes distributed across repository evolution remains poorly understood. We introduce PRWeaver, a benchmark of 208 execution-validated attacks from ten real-world repositories, each instantiated under four matched review renderings (832 renderings in total). We evaluate three PR-auditing agents across six auditor-model systems. Across all systems, decomposing an attack changes detection by at most five percentage points, showing that commit boundaries alone do not explain evasion. In contrast, per-PR interleaving at $N=16$ and coherent carrier fusion reduce detection by 5-13 and 10-18 points, respectively. Under whole-window review at $N=24$, detection falls to 16-22%, compared with 50-60% under per-PR review. These results show that access to repository history is insufficient: concealment becomes most effective when benign and malicious changes jointly occupy the auditor's active review context or when the stated purpose plausibly accounts for the attack-bearing diff.