arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Antaeus: 通过上下文锚定的LLM推理发现仓库级逻辑漏洞

Antaeus: Hunting Repository-Level Logic Vulnerabilities via Context-Grounded LLM Reasoning

Michele Armillotta, Nicolò Romandini, Rebecca Montanari, Lorenzo Cavallaro

arXiv 2607.01138首次发表:更新:

发表机构

University College London; University of Bologna(伦敦大学学院; 博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Antaeus框架,通过仓库级代码上下文锚定LLM推理,结合函数优先级排序、上下文推理、比较验证和结构化报告,有效检测逻辑漏洞,在28个仓库中检测出15个漏洞,优于基线方法。

AI 中文摘要

基于LLM的漏洞检测器在识别内存安全漏洞和可通过既定安全属性表达的漏洞类别方面显示出有希望的结果。然而,逻辑漏洞提出了不同的挑战,因为它们的识别需要推断应用特定的安全不变性和关于预期行为的隐含假设。即使是最前沿的智能体模型也难以应对,因为这些不变性通常是隐含的,并埋藏在无关代码中。受此差距的驱动,我们提出了Antaeus,一个将LLM推理锚定在仓库级代码上下文中的逻辑漏洞检测框架。Antaeus遵循一个仓库规模的流水线,结合函数优先级排序、上下文锚定推理、比较验证和结构化报告。它使用轻量级的仓库范围安全信号对函数进行排序,将昂贵的LLM分析引导到相关代码上,减少调用、成本和分类工作。对于每个优先排序的函数,Antaeus将本地代码上下文与应用程序功能、安全资源和信任边界的仓库级视图相结合。这使得能够推理该函数在更广泛的应用程序中如何执行,而不是作为一个孤立的片段。Antaeus识别安全敏感汇点,推导安全执行的条件,并检查这些条件是否在本地满足。候选发现经过比较验证,修剪反映项目范围规范而非独特违规的问题。最后,Antaeus报告汇点、违反的安全条件和证据,使发现可操作且可追溯。我们在28个已确认逻辑漏洞的仓库上评估Antaeus,并将其与函数级和智能体模型进行比较。Antaeus检测并解释了15个漏洞,在可比的令牌使用和成本下优于基线方法。

英文摘要

LLM-based vulnerability detectors have shown promising results in identifying memory-safety bugs and vulnerability classes whose violations can often be expressed through established security properties. Logic vulnerabilities, however, pose a different challenge, as their identification requires inferring application-specific security invariants and implicit assumptions about intended behavior. Even frontier agentic models struggle because these invariants are often implicit and buried among unrelated code. Motivated by this gap, we present Antaeus, a framework for detecting logic vulnerabilities that grounds LLM reasoning in repository-level code context. Antaeus follows a repository-scale pipeline combining function prioritization, context-grounded reasoning, comparative validation, and structured reporting. It ranks functions using lightweight repo-wide security signals, directing costly LLM analysis toward relevant code and reducing calls, cost, and triage effort. For each prioritized function, Antaeus combines local code context with a repository-level view of the application's functionality, security resources, and trust boundaries. This enables reasoning about how the function is executed within the broader application rather than as an isolated snippet. Antaeus identifies security-sensitive sinks, derives safety conditions for safe execution, and checks whether they are locally satisfied. Candidate findings undergo comparative validation, pruning concerns that reflect project-wide norms rather than distinctive violations. Finally, Antaeus reports sinks, violated safety conditions, and evidence, making findings actionable and traceable. We evaluate Antaeus on 35 repositories with confirmed logic vulnerabilities and compare it against function-level and agentic models. Antaeus detects and explains 20 vulnerabilities, outperforming baselines with comparable token usage and cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑