arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DREA:用于仓库级漏洞检测的解耦推理与探索代理

DREA: Decoupled Reasoning and Exploration Agents for Repository-Level Vulnerability Detection

Mingyang Sun, Guozhu Meng

arXiv 2607.13439首次发表:更新:

AI 中文总结

研究针对大语言模型用于仓库级漏洞检测时的不足,提出DREA框架,通过规划与探索代理解耦,实现目标导向上下文获取。构建RepoPairBench评估,提升配对正确性,降低成本,还发现安全推理质量是当前大语言模型的瓶颈。

AI 中文摘要

大语言模型因其强大的代码理解能力越来越多地应用于漏洞检测,但现有方法大多依赖固定程序分析规则提取的孤立函数或上下文。当漏洞跨越多个函数或文件时,无法自适应探索仓库级依赖以收集足够上下文,影响检测可靠性。我们提出DREA,一个用于仓库级漏洞检测的假设驱动框架。它通过两个协作代理将推理与探索解耦:由先进大语言模型支持的规划代理形成漏洞假设并指导调查,由轻量级模型驱动的探索代理按需检索仓库级上下文。目标导向的上下文获取是检测改进的主要来源,将大量令牌的探索卸载到本地模型使推理经济可行。为支持评估,我们构建了RepoPairBench,一个基于真实世界项目中经过验证的Python漏洞修复对的仓库基准。除了二进制检测准确率,我们引入推理正确性评估来评估模型的理由是否与记录的漏洞机制匹配。在三个大语言模型上,DREA将配对正确性从19 - 26%提高到30 - 42%,同时将超过93%的令牌卸载到探索器,将估计的可计费API成本降低16 - 48倍。推理正确性分析进一步表明,DREA和仅函数基线的26 - 55%的真阳性是由有缺陷的理由支持的正确预测,确定安全推理质量是当前大语言模型的共同瓶颈。

英文摘要

Large language models (LLMs) are increasingly applied to vulnerability detection due to their strong code comprehension capabilities, but most existing approaches rely on isolated functions or context extracted by fixed program-analysis rules. These methods cannot adaptively explore repository-level dependencies to gather sufficient context when vulnerabilities span multiple functions or files, compromising detection reliability. We present DREA (Decoupled Reasoning and Exploration Agents), a hypothesis-driven framework for repository-level vulnerability detection. DREA decouples reasoning from exploration through two collaborating agents: a planning agent backed by an advanced LLM that forms vulnerability hypotheses and directs the investigation, and an explorer agent powered by a lightweight model that retrieves repository-level context on demand. Goal-directed context acquisition is the primary source of detection improvement in this design, while offloading token-heavy exploration to the local model keeps inference economically tractable. To support evaluation, we construct RepoPairBench, a repository-grounded benchmark of validated Python vulnerability-fix pairs from real-world projects. Beyond binary detection accuracy, we introduce a reasoning correctness evaluation to assess whether a model's rationale matches the documented vulnerability mechanism. Across three LLMs, DREA improves Pair-Correctness from 19-26% to 30-42% while offloading over 93% of tokens to the explorer, reducing estimated billable API cost by a factor of 16-48. Reasoning correctness analysis further reveals that 26-55% of true positives, for both DREA and the function-only baseline, are correct predictions supported by flawed rationales, identifying security reasoning quality as a shared bottleneck for current LLMs.

CommentsAccepted for presentation at Internetware 2026. 12 pages, 4 figures, and 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑