发表机构
Pazhou Laboratory(琶洲实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体持久内存的谄媚与泄漏问题,提出FC、PS两种推理时准入设计,实验显示其可降低失败率与跨域泄漏,但需解决有益内存使用保留问题。
AI 中文摘要
持久内存可提升LLM智能体的个性化能力,但也会引发谄媚行为与跨域信息泄漏问题。我们区分两类管控决策:准入(admission),用于确定哪些召回信息会进入工作上下文;呈现(presentation),用于确定准入信息的表达方式。我们实现了两种无需重新训练的推理时设计:因子编译准入(FC),用于评估完整的内存条目;权限语义准入(PS),用于将条目分解为带类型的单元;两者均通过确定性策略将裁决属性转化为资格决策。我们在包含四个主干模型的开发套件及一个外部基准上进行评估,外部基准包含四个任务,每个任务有300个样本。与逐字注入相比,FC和PS在外部基准上降低了经评判者评估的汇总失败率,分别减少6.7和8.8个百分点(p值分别为2.7e-7和4.1e-12),开发集的跨域泄漏最多降低29.5个百分点。查询条件门控基线在客观事实失败或汇总失败上无显著变化。在匹配的准入预算下,经Holm校正后,PS在外部客观事实判断上优于随机选择和基于相关性的选择。在固定呈现方式的情况下,收紧准入可进一步将跨域失败降低17.5个百分点(p值为1.6e-4);相比之下,相同裁决输出的两种呈现方式之间的比较未通过多重比较校正。两种设计均增加了个性化失败,且PS未达到预先注册的改进及个性化保留标准。这些结果支持将准入与呈现分开评估:选择质量可带来特定任务的安全提升,而保留有益的内存使用仍未解决。
英文摘要
Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We implement two inference-time designs without retraining: factor-compiled admission (FC), which assesses whole memory entries, and permission-semantic admission (PS), which decomposes entries into typed units; both translate adjudicated attributes into eligibility decisions via deterministic policies. We evaluate on a four-backbone development suite and an external benchmark with four tasks of 300 samples each. Relative to verbatim injection, FC and PS reduce pooled judge-assessed failure rates on the external benchmark by 6.7 and 8.8 percentage points (p = 2.7e-7 and 4.1e-12), and development-set cross-domain leakage falls by up to 29.5 percentage points. A query-conditioned gating baseline shows no significant change in objective-fact failure or pooled failure. Under matched admission budgets, PS outperforms random and relevance-based selection on external objective-fact judgment after Holm correction. Holding presentation fixed, tightening admission cuts cross-domain failure by a further 17.5 percentage points (p = 1.6e-4); in contrast, no comparison between two renderings of identical adjudicated outputs survives multiple-comparison correction. Both designs increase personalization failures, and PS misses the preregistered improvement and personalization-preservation criteria. These results support evaluating admission and presentation separately: selection quality provides task-specific safety gains, while preserving beneficial memory use remains unresolved.
CommentsUnder review at AAMAS 2027