我的助手会记得我的过敏症吗?当对话记忆被压缩时,个人LLM助手会遗忘什么
Will My Assistant Remember My Allergy? What Personal LLM Assistants Forget When Conversation Memory Is Compressed
浏览论文内容
中文总结 AI 辅助
研究发现,个人LLM助手在压缩对话记忆时,安全关键事实(如过敏信息)的保留率从97%骤降至0-1%,表明压缩缓存仅适用于推理复用,需额外可审计存储与主动询问界面。
中文摘要 AI 辅助
个人LLM助手(健康伴侣、老年护理代理、无障碍辅助工具)的评判标准在于它们对一个人的记忆:例如,一句随口提及、几天后需要的药物或过敏信息。隐私需求促使它们部署在设备端,而一个月的对话量可能超过模型自身的权重规模,因此必须采用驱逐策略来决定缓存遗忘哪些内容。基准测试报告称,驱逐策略在20%的预算下能保留此类事实,但这些测试压缩的是已经包含用户未来问题的提示词,这种预见性是没有缓存复用助手所不具备的。若在压缩之后才揭示问题,这一优势便荡然无存:在我们构建的PA-Bench上的100个助手对话中,一句随口提及的过敏信息能存活到需要它的那个问题的概率仅为0%至1%,而完整记忆下的存活率为97%。其原因在于预算而非评分器:我们评估的所有无训练策略都未能将该事实排名足够靠前,而能够保留该事实所需的预算又大到不值得进行压缩。压缩缓存是一种推理复用机制,而非持久化层:安全关键事实需要一个可审计的情节存储作为补充,以及一个主动询问而非凭空编造的界面。
英文摘要
Personal LLM assistants (health companions, elder-care agents, accessibility aides) are judged by what they remember about a person: a medication or an allergy mentioned in passing and needed days later. Privacy pushes them on-device, where a month of conversation can outgrow the model's own weights, so an eviction policy must decide what the cache forgets. Benchmarks report that eviction keeps such facts at a 20% budget, but they compress a prompt that already contains the user's future question, foresight no cache-reusing assistant has. Hide the question until after compression and the advantage vanishes: on PA-Bench, 100 assistant conversations we construct, an allergy mentioned in passing survives to the question that needs it 0--1% of the time, against 97% with full memory. The cause is the budget, not the scorer: none of the training-free policies we evaluate ranks the fact high enough, and the budget that would keep it is too large to bother compressing. A compressed cache is an inference-reuse mechanism, not a persistence layer: safety-critical facts need an auditable episodic store alongside it, and an interface that asks rather than invents.
发表机构
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。