AI 中文总结
CAPTURE通过神经微分方程信念追踪器等技术,在个性化LLM智能体中区分偏好漂移与内存投毒,提升胜率并限制投毒成功率,验证了偏好真实性建模的价值。
AI 中文摘要
个性化语言智能体利用持久内存随时间适配用户,但该机制也形成了攻击面。当新信息与存储的偏好冲突时,智能体必须区分真实偏好漂移与临时上下文偏移、歧义或对抗性内存投毒。我们将该问题表述为关于潜在用户状态的连续时间部分可观测决策过程,并证明仅基于新近度和来源的规则为何不足。CAPTURE通过神经微分方程信念追踪器、多时间尺度内存账本、不确定性触发澄清及引用内存的反事实审计解决此类歧义。在96个用户的480个保留情节上,CAPTURE的胜率达71.5%,而相同监督的基线为69.3%,最强启发式基线为66.1%;其将固定策略投毒成功率限制在11.5%,同时接受83.5%的真实偏好更新。在可访问发布权重的自适应攻击者下,攻击成功率升至24.7%,暴露了真实的适配-安全权衡。我们还在独立构建的基准上对冻结系统进行零样本评估,并重放来自40个用户、时长2至3周的纵向交互历史,结果表明显式建模偏好真实性可提升内存增强大语言模型智能体的个性化与鲁棒性。
英文摘要
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
CommentsUnder review at ICLR 2027