arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SelfOp:一种用于自我改进安全智能体的优化算法

SelfOp: An Optimization Algorithm for Self-Improving Security Agents

Saad Ullah, Yigitcan Kaya, Christopher Kruegel, Giovanni Vigna, Gianluca Stringhini

arXiv 2609.22792首次发表:更新:

发表机构

Boston University; UC Santa Barbara(波士顿大学; 加州大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SelfOp提出一种受链式法则启发的文本梯度下降算法,自动优化安全智能体的任务上下文,无需修改模型权重,在CyberGym基准上以少于200个示例实现显著自我改进,且技能可跨模型迁移。

AI 中文摘要

LLM智能体越来越多地被用于安全任务:漏洞发现、漏洞利用复现和补丁生成。在模型层面改进它们需要专家示范或可计算的奖励,而安全任务很少能提供这些:轨迹成本高昂,失败难以诊断,奖励稀疏且不可计算。因此,工作转向了工具框架(harness)和上下文,但手动调整需要特定任务的专业知识且扩展性差,而自动化方法则依赖稀缺的真实标签、更强的优化器模型,或未经引导的提出-评估循环,这些循环退化为昂贵的试错。我们引入了SelfOp,一种自动改进冻结安全智能体任务上下文(指令、技能和参考文档)的算法,无需修改其执行工具框架和模型权重。SelfOp将上下文优化视为受链式法则启发的文本梯度下降:从单个实例的结果出发,它通过评估器、智能体的轨迹以及塑造其行为的上下文工件向后传播错误信号,生成每个实例的文本梯度。梯度通过聚类、排序和过滤在实例间累积,并且仅在跨实例共识下提交。一个收敛检测器监控梯度信号本身,并在上下文吸收了训练数据中的可泛化信息后停止,无需保留验证数据。我们在CyberGym(一个真实世界漏洞复现任务的基准)上评估了SelfOp。在少于200个训练示例的情况下,SelfOp为GPT-5.4-mini(使用Codex)带来了17个百分点的自我改进,足以超过前沿GPT-5.4基线6个百分点,并为GPT-5.4本身带来了18.5个百分点的自我改进。优化后的技能还能跨模型迁移,这突显了SelfOp优化的技能学习的是可泛化的任务知识,而非模型特定的模式。

英文摘要

LLM agents are increasingly used for security tasks: vulnerability discovery, exploit reproduction, and patch generation. Improving them at the model level demands expert demonstrations or computable rewards, which security tasks rarely offer: traces are costly, failures hard to diagnose, rewards sparse, and non-computable. Efforts thus shift to the harness and context, but manual tuning needs task-specific expertise and scales poorly, while automated methods rely on scarce ground truth, stronger optimizer models, or unguided propose-and-evaluate loops that reduce to costly trial and error. We introduce SelfOp, an algorithm that automatically improves a frozen security agent's task context (instructions, skills, and reference documents), without modifying its execution harness and model weights. SelfOp casts context optimization as chain-rule-inspired textual gradient descent: from a single instance's outcome, it propagates error signals backward through the evaluator, the agent's trajectory, and the context artifacts that shaped its behavior, yielding per-instance textual gradients. Gradients are accumulated across instances by clustering, ranking, and filtering, and committed only under cross-instance consensus. A convergence detector monitors the gradient signal itself and stops once the context has absorbed the generalizable information in the training data, without held-out validation data. We evaluate SelfOp on CyberGym, a benchmark of real-world vulnerability reproduction tasks. With fewer than 200 training examples, SelfOp yields a 17-point self-improvement for GPT-5.4-mini (with Codex), enough to surpass the frontier GPT-5.4 baseline by 6 points, and an 18.5-point self-improvement for GPT-5.4 itself. The optimized skills also transfer across models, highlighting that SelfOp-optimized skills learn generalizable task knowledge not model-specific patterns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑