arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

假设成本低,验证成本高:智能体漏洞发现的防御面

Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery

Kaikai Zhang, Zihan Zhang, Yuchong Xie, Zesen Liu, Shuangjie Yao, Zhixiang Zhang, Dongdong She

arXiv 2609.35909首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自主LLM智能体漏洞发现中的假设-验证不对称性,提出RedHerring方法,通过插入可证明安全的诱饵转移验证资源,在33个项目中减少38.7%-60.4%的真实漏洞发现。

AI 中文摘要

自主LLM智能体将漏洞发现转变为仓库规模的搜索:它们生成大量漏洞假设,但在有限预算下只能验证其中一部分。我们表明,自主漏洞发现存在假设-验证不对称性,其中通过可达性分析、执行和概念验证构建来验证候选假设的成本远高于形成假设。在有限资源预算下,这使得自主发现成为一个资源受限的选择性验证过程,进一步将验证工作暴露为独特的防御面。我们提出RedHerring,它插入可证明安全的诱饵,将验证工作从真实漏洞中转移。每个诱饵结合了吸引验证的CVE衍生漏洞链和保持其危险汇点不可达的虚假桥接。私有证书使防御者能够高效验证此属性,而从发布的仓库中确立相同事实则需要解决计算上困难的问题。RedHerring进一步将每个诱饵适配到目标仓库,使其看起来像普通程序逻辑。在33个OSS-Fuzz项目、70个评估实例和五个模型在匹配预算下的实验中,RedHerring将发现的真实漏洞减少了38.7%-60.4%。轨迹分析显示,智能体花费30.6%-51.5%的完成令牌和估计32.5%-49.9%的运行时间验证诱饵,表明RedHerring将固定搜索预算的很大一部分重定向到诱饵上。当明确告知可能存在诱饵时,智能体会调整其搜索策略,但RedHerring相对于知情基线仍将发现的漏洞减少了37.2%,表明其有效性不依赖于诱饵的保密性。

英文摘要

Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it. Under a finite resource budget, this makes autonomous discovery a resource-bounded selective-verification process, further exposing verification effort as a unique defense surface. We present RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities. Each decoy combines a CVE-derived vulnerability chain that attracts verification with a false bridge that keeps its dangerous sink unreachable. A private certificate lets the defender verify this property efficiently, while establishing the same fact from the released repository requires solving a computationally hard problem. RedHerring further adapts each decoy to the target repository so that it reads as ordinary program logic. Across 33 OSS-Fuzz projects, 70 evaluation instances, and five models under matched budgets, RedHerring reduces real vulnerabilities discovered by 38.7-60.4%. Trajectory analysis shows that agents spend 30.6-51.5% of completion tokens and an estimated 32.5-49.9% of runtime verifying decoys, showing that RedHerring redirects a substantial fraction of the fixed search budget toward decoys. When explicitly informed that decoys may be present, the agent adapts its search strategy, yet RedHerring still reduces vulnerabilities discovered by 37.2% relative to an informed Baseline, showing that its effectiveness does not depend on decoy secrecy.

Comments37 pages. Project page: https://xxbai.space/redherring/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑