发表机构
UIUC; Purdue University(伊利诺伊大学厄巴纳-香槟分校; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出诊断框架Reveree,从三层级评估LLM逆向工程智能体,发现基础模型主导性能、失败集中于理解阶段、大/新/贵模型未必更强,且发布该框架供社区使用。
AI 中文摘要
逆向工程(RE)对于恶意软件分析、漏洞发现等安全任务至关重要,而大型语言模型(LLM)智能体正越来越多地具备自主执行逆向工程的能力。夺旗赛(CTF)逆向工程挑战已成为衡量该能力的标准替代方案,但评估仅基于单一标准:智能体是否成功夺旗。这种成功率既无法揭示智能体在逆向工程流程的哪个环节失败,也无法判断成功是源于对二进制文件的分析还是对公开解决方案的回忆。本文中,我们提出Reveree这一诊断框架,它从三个层级对LLM逆向工程智能体的轨迹进行评分:成功率、通过八阶段逆向工程架构的里程碑进度,以及其行为的行为特征。理解阶段由经人类专家验证的、不依赖结果的LLM评判器评分,其余所有阶段均通过确定性方式验证。我们使用Reveree评估了9个前沿模型和4种提示策略,涉及88个picoCTF和NYU-CTF挑战。研究发现,基础模型主导性能表现,而提示策略是次要的、依赖于模型的影响因素;令人惊讶的是,更大、更新或更昂贵的模型并非始终更强。我们还发现,失败集中在逆向工程流程的理解阶段,额外的预算、坚持度或推理努力仅能挽救少数失败,这指向能力限制而非资源限制。关于记忆,模型可仅通过挑战描述复现picoCTF的旗标,但NYU-CTF显示可测量的记忆极少,且多数成功可在表面扰动下保持,表明真正的分析与记忆共存。我们向社区发布Reveree。
英文摘要
Reverse engineering (RE) is critical to security tasks such as malware analysis and vulnerability discovery, and large language model (LLM) agents are increasingly able to perform it autonomously. Capture-the-flag (CTF) RE challenges have become the standard proxy for measuring this capability, but evaluation rests on a single criterion: whether the agent captures the flag. This solve rate reveals neither where in the RE process an agent fails nor whether a success reflects analysis of the binary or recall of a public solution. In this paper, we propose Reveree, a diagnostic framework that scores an LLM RE agent's trajectory at three tiers: solve rate, milestone progress through an eight-stage RE schema, and a behavioral profile of its actions. Comprehension stages are scored by an outcome-blinded LLM judge validated against a human expert; all other stages are verified deterministically. Using Reveree, we evaluate nine frontier models and four prompting strategies on 88 picoCTF and NYU-CTF challenges. We find that the base model dominates performance, whereas prompting strategy is a secondary, model-dependent effect. Surprisingly, larger, newer, or costlier models are not reliably stronger. We also find that failures concentrate at the comprehension stages of the RE process, and that extra budget, persistence, or reasoning effort rescues few of them, pointing to a competence limit rather than a resource limit. Regarding memorization, while models reproduce picoCTF flags from challenge descriptions alone, NYU-CTF shows minimal measurable recall, and most solves survive surface perturbation, indicating that genuine analysis coexists with memorization. We release Reveree to the community.