arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13734cs.CLcs.AI

PolicyMem:面向LLM治理的几何策略记忆

PolicyMem: Geometric Policy Memory for LLM Governance

Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao, Haoyu Wang, Hanghang Tong, Haifeng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

PolicyMem提出几何策略记忆,将自然语言策略外部化为低秩子空间对象,通过投影能量实现检测-重写-验证循环,在五个基准上达到最先进的不安全行为检测并支持策略归因与验证。

中文摘要 AI 辅助

随着大语言模型(LLMs)越来越多地部署于现实世界的高风险应用中,有效的治理变得至关重要。现有的安全防护措施主要遵循两种范式:基于学习的防护提供强大的语义判别能力,但将策略行为与训练模型和分类体系耦合在一起;而可编程框架提供灵活的控制,但需要大量的手动提示词和工作流工程。这两种方式都没有将策略外部化为可复用的操作状态,使得在检测、干预和验证环节中一致地复用策略证据变得困难。在本文中,我们提出了PolicyMem,一种几何策略记忆,它将自然语言策略外部化为可复用的几何记忆对象,这些对象由共享表示空间中的低秩子空间表示。一个记忆写入器将自然语言策略编译为策略记忆槽,查询-响应对通过投影能量读取策略记忆。由此产生的策略证据档案直接调节安全判定,并被复用于策略归因和干预后验证。结合响应重写器,PolicyMem实现了用于LLM治理的检测-重写-验证循环。在五个广泛使用的基准测试中,PolicyMem在实现不安全行为检测的最先进性能的同时,通过共享策略记忆实现了有效的策略归因、重写和干预后验证。

英文摘要

As large language models (LLMs) are increasingly deployed in real-world high-stakes applications, effective governance has become essential. Existing safeguards largely follow two paradigms: learning-based guards provide strong semantic discrimination but couple policy behavior to trained models and taxonomies, while programmable frameworks offer flexible control but require substantial manual prompt and workflow engineering. Neither externalizes policies as reusable operational states, making it difficult to consistently reuse policy evidence across detection, intervention, and verification. In this paper, we introduce PolicyMem, a geometric policy memory that externalizes natural-language policies as reusable geometric memory objects represented by low-rank subspaces in a shared representation space. A memory writer compiles natural-language policies into policy memory slots, and query-response pairs read the policy memory through projection energy. The resulting policy-evidence profile directly mediates the safety verdict and is reused for policy attribution and post-intervention verification. Coupled with a response rewriter, PolicyMem enables a detect-rewrite-verify loop for LLM governance. Across five widely used benchmarks, PolicyMem achieves state-of-the-art unsafe behavior detection while enabling effective policy attribution, rewriting, and post-intervention verification through the shared policy memory.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • NEC Laboratories America(NEC美国实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑