arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11469cs.CRcs.AIcs.SE

智能体网络安全的下一个挑战:一个现实、无污染的逆向工程基准

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

Jeremy Spence, Nicholas Assaderaghi, Feng Xiao, Jinhao Zhu, Nikil Ravi, Xiangyu Qi, Matthew Jagielski, Raluca Ada Popa, Eric Wallace, Guannan Wei, Yangruibo Ding, Zhuo Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对智能体逆向工程基准的缺陷,构建SRE-Bench基准,评估五款前沿LLMs的逆向工程能力,发现RE仍未解决,强调其为智能体网络安全重要前沿及SRE-Bench的测试价值。

中文摘要 AI 辅助

当源代码可供分析时,AI智能体的网络安全能力正在快速提升,但对网络安全至关重要的大部分软件(包括恶意软件、固件和专有应用程序)仅以二进制形式提供。分析此类软件需要逆向工程(RE):在进行有意义的分析之前恢复程序语义。然而,评估智能体RE面临一个根本性挑战:基准实例必须在LLMs的训练数据中作为源代码不可见,以防止模型通过识别它们而非真正分析来走捷径,同时还要匹配真实软件的规模和反分析保护措施。遗憾的是,现有基准无法同时满足这些要求。为此,我们推出SRE-Bench,这是首个现实、无污染的RE基准。SRE-Bench由RE专家从零开始构建,耗时超过5000小时,包含19个私人的、真实世界规模的程序,平均代码行数为16.9K。我们还开发了44种内部反分析原语,生成262个二进制实例和1572个确定性分级任务。我们对五个前沿LLMs(GPT-5.6-sol、Claude-Opus-5、GPT-5.5、Grok-4.5和GLM-5.2)的评估显示,RE在很大程度上仍未解决:最强的模型GPT-5.6-sol的每个实例得分为61.4%,且仅完全解决31.5%的实例。我们的分析进一步揭示,智能体的行为与人类工程师不同,智能体对编译器优化和静态链接相对不敏感。受控消融实验还证实,污染控制和现实规模都是必不可少的。这些结果表明,强大的源代码安全能力尚未转移到二进制分析,凸显RE是智能体网络安全的重要前沿,而SRE-Bench是衡量进展的严格测试平台。

英文摘要

AI agents are rapidly improving in cybersecurity when source code is available, yet much of the software most consequential to security, including malware, firmware, and proprietary applications, exists only as binaries. Analyzing such software requires reverse engineering (RE): recovering program semantics before analysis can proceed. Evaluating agentic RE poses a fundamental challenge: realistic benchmark instances must (1) be absent from LLMs' training data to prevent shortcuts by memorization, and (2) reflect the scale and anti-analysis protections of real-world binaries. We introduce SRE-Bench, the first realistic, contamination-free RE benchmark. Built from scratch by RE experts with over 5,000 expert hours, SRE-Bench comprises 19 private, real-world-scale programs averaging 16.9K lines of code. We further developed 44 in-house anti-analysis mechanisms, yielding 262 binary instances and 1,572 deterministically graded tasks. We evaluated 13 agentic settings across 11 models: eight in public-facing settings and five in internal unconstrained settings with cyber safeguards disabled and no budget cap. Realistic RE remains challenging for frontier agents: GPT-5.6-Sol and Claude-Fable-5.1, despite strong source-code security capabilities, fully solve only 31.5% and 26.9% of graded instances, suggesting that success in source-code security does not translate into effective binary analysis. Without a budget cap and safety guard, GPT-6-Astra achieves a near-perfect pass@4 score, yet reliably identifying the correct candidate remains difficult. Agents are largely insensitive to compiler optimization and static linking, and ablations confirm that both contamination control and realistic scale are essential to understanding agents' RE capability. These findings highlight RE as a distinct frontier for agentic cybersecurity and establish SRE-Bench as a rigorous testbed for measuring progress.

发表机构

  • UC Berkeley(加州大学伯克利分校)
  • Vals AI(瓦尔兹人工智能公司)
  • Tufts University(塔夫茨大学)
  • UCLA(加州大学洛杉矶分校)
  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑