arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07403cs.AI

深度防御大语言模型:评估记忆门控对抗激活诱导与记忆诱导的谄媚行为

Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy

Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM长期记忆引发的谄媚行为,提出内外分离的深度防御框架,验证记忆门控过滤有效,而反向激活引导无显著额外增益。

中文摘要 AI 辅助

长期记忆使大型语言模型(LLMs)能够在交互中维持个性化上下文,但检索到的用户历史可能引发记忆诱导的谄媚行为,导致模型倾向于存储的用户信念而非客观证据。现有防御措施主要作用于检索到的上下文,且很少与内部行为偏差联合评估。我们引入一个$2 \times 2$的深度防御框架,将内部激活引导与外部记忆处理分开。我们从100对提示中提取谄媚引导方向,并在MemSyco-Bench上评估四个开放权重模型,涵盖10个引导系数和五种记忆防御配置(所有1,550个条目的答案;防御条件在固定的250条目子样本上评判),使用三个LLM评判者。五种配置中有三种是新的(重写每条记忆、路由器门控(保留、重写或丢弃每条记忆)以及丢弃所有记忆);另外两种是MemSyco的基线。选择性路由器门控过滤比完全移除记忆保留了显著更多的MemSyco平均准确率,并且当模型被引导向谄媚时,这种分离仍然存在。在Llama 3.1 8B上使用路由器门控时,轻度反向引导($\alpha = -1.5$)将评判者平均谄媚率从35.80%降至31.32%,而平均准确率从43.99%变为43.31%;这种降低在所有三个评判者下方向一致,但统计上不显著(配对$p = 0.08$至$0.63$,基于149个条目)。外部记忆过滤是设计中经得起考验的部分;我们的数据并未显示反向引导对其有额外贡献。

英文摘要

Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a $2 \times 2$ defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients and five memory-defense configurations on MemSyco-Bench (answers for all 1,550 items; defense conditions judged on a fixed 250-item subsample), with three LLM judges. Three of the five configurations are new (rewriting every memory, a Router Gate that keeps, rewrites, or drops each memory, and dropping all memory); the other two are MemSyco's baselines. Selective Router Gate filtering preserves substantially more of MemSyco's average accuracy than complete memory removal, and this separation persists when the models are steered toward sycophancy. On Llama 3.1 8B with Router Gate, mild inverse steering ($α= -1.5$) lowers judge-averaged sycophancy from 35.80% to 31.32% while average accuracy moves from 43.99% to 43.31%; this reduction has the same direction under all three judges but is not statistically significant (paired $p = 0.08$ to $0.63$ on 149 items). External memory filtering is the part of the design that holds up; our data do not show that inverse steering adds to it.

发表机构

  • Urbana CS Club
  • CaML

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑