arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mastermind: 基于策略学习的仓库级漏洞复现

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

Mingzhe Du, Luu Anh Tuan, Tianyi Wu, Renyang Liu, Zhijiang Guo, Dong Huang, See-Kiong Ng

arXiv 2607.01764首次发表:更新:

发表机构

National University of Singapore; Nanyang Technological University; The Hong Kong University of Science and Technology (Guangzhou)(新加坡国立大学; 南洋理工大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Mastermind双循环框架,通过策略学习(SFT+GRPO)提升LLM代理在仓库级漏洞复现中的表现,在260个训练任务和200个测试任务上达到84.5%通过率,并成功迁移至不同执行器。

AI 中文摘要

仓库级漏洞复现是一项高要求的软件工程任务:代理必须检查代码库,推断到达漏洞路径的输入语法,构建概念验证(PoC),并验证崩溃在修补后的构建中消失。最近的LLM代理在方法正确时通常能执行这些步骤,但仍因选择错误策略而失败。本文认为,策略而非完整动作轨迹是此类SE代理的正确学习单元:它足够紧凑以优化,足够具体以指导执行,且足够稳定以跨尝试存储和重用。我们提出Mastermind,一个双循环框架,将可迁移的策略学习与任务特定经验分离。可训练的计划器通过SFT和基于里程碑的GRPO学习可重用的漏洞复现策略,而经验循环维护任务本地策略记录以指导后续尝试。计划器独立于执行器训练,使得策略学习能够改进多个冻结执行器而不修改其动作生成能力。我们在CyberGym上使用260个训练任务和200个保留评估任务评估Mastermind。以GPT-5.5作为冻结执行器,Mastermind达到84.5%的通过率,优于开放书PoC上下文(60.0%)、Best-of-8采样(63.0%)和迭代改进(77.0%)。相同的计划器还将GPT-5.4 mini和GLM~5.1从45.0%和58.5%提升至60.0%和71.0%。这些结果表明,学习高层策略是改进仓库级SE代理的有效且可迁移的机制。

英文摘要

Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path, construct a proof-of-conceptv(PoC), and verify that the crash disappears on the patched build. Recent LLM agents can often execute these steps when the approach is correct, yet they still fail by choosing the wrong strategy. This paper argues that strategy, rather than the full action trajectory, is the right learning unit for such SE agents: it is compact enough to optimize, concrete enough to guide execution, and stable enough to store and reuse across attempts. We present Mastermind, a dual-loop framework that separates transferable strategy learning from task-specific experience. A trainable planner learns reusable vulnerability-reproduction strategies through SFT and milestone-based GRPO, while an experience loop maintains task-local strategy records that guide subsequent attempts. The planner is trained independently of the executor, allowing strategy learning to improve multiple frozen executors without modifying their action-generation capability. We evaluate Mastermind on CyberGym using 260 training tasks and 200 held-out evaluation tasks. With GPT-5.5 as the frozen executor, Mastermind achieves an 84.5% pass rate, outperforming open-book PoC context (60.0%), Best-of-8 sampling (63.0%), and iterative improvement (77.0%). The same planner also improves GPT-5.4 mini and GLM~5.1 from 45.0% and 58.5% to 60.0% and 71.0%. These results demonstrate that learning high-level strategies is an effective and transferable mechanism for improving repository-scale SE agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑