通过可验证奖励、经验合成和持续记忆对代理LLM进行审计化技能图自改进
Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory
- OWASP Fairfax, VA, USA
- Kleiner Perkins
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ASG-SI框架,通过可验证奖励、经验合成和持续记忆实现代理LLM的审计化技能图自改进,以提升安全性和可追溯性。
AI中文摘要:
强化学习正被越来越多地用于将大语言模型转变为能够长期行动、调用工具并管理记忆的代理系统。尽管最近的研究通过工具学习、可验证奖励和持续训练展示了性能提升,但部署的自我改进代理却面临未解决的安全性和治理挑战:优化压力可能激励奖励黑客,行为漂移难以审计或重现,改进往往纠缠在不透明的参数更新中,而非可重用的、可验证的成果。本文提出了审计化技能图自改进(ASG-SI)框架,将自我改进视为代理迭代编译为不断增长、可审计的技能图的过程。每个候选改进均从成功的轨迹中提取,归一化为具有显式接口的技能,并在通过验证者支持的回放和合同检查后才被提升。奖励被分解为可重构的组成部分,这些组成部分源于可回放的证据,使推广决策和学习信号的独立审计成为可能。ASG-SI进一步整合了经验合成以实现可扩展的压力测试和持续记忆控制,以在受限上下文中保持长周期性能。我们展示了完整的系统架构、威胁模型和安全分析,并提供了一个完全可运行的参考实现,展示了验证者支持的奖励构建、技能编译、审计日志记录以及在持续任务流下的可测量改进。ASG-SI将代理自我改进重新定义为可验证、可重用能力的积累,为自我改进AI代理的可重复评估和操作治理提供了一条实用路径。
英文摘要:
Reinforcement learning is increasingly used to transform large language models into agentic systems that act over long horizons, invoke tools, and manage memory under partial observability. While recent work has demonstrated performance gains through tool learning, verifiable rewards, and continual training, deployed self-improving agents raise unresolved security and governance challenges: optimization pressure can incentivize reward hacking, behavioral drift is difficult to audit or reproduce, and improvements are often entangled in opaque parameter updates rather than reusable, verifiable artifacts. This paper proposes Audited Skill-Graph Self-Improvement (ASG-SI), a framework that treats self-improvement as iterative compilation of an agent into a growing, auditable skill graph. Each candidate improvement is extracted from successful trajectories, normalized into a skill with an explicit interface, and promoted only after passing verifier-backed replay and contract checks. Rewards are decomposed into reconstructible components derived from replayable evidence, enabling independent audit of promotion decisions and learning signals. ASG-SI further integrates experience synthesis for scalable stress testing and continual memory control to preserve long-horizon performance under bounded context. We present a complete system architecture, threat model, and security analysis, and provide a fully runnable reference implementation that demonstrates verifier-backed reward construction, skill compilation, audit logging, and measurable improvement under continual task streams. ASG-SI reframes agentic self-improvement as accumulation of verifiable, reusable capabilities, offering a practical path toward reproducible evaluation and operational governance of self-improving AI agents.