arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03130cs.CRcs.CLcs.LG

DP-MemView:面向长期大语言模型智能体的属性级对话记录隐私的内存接口

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

Jong Wook Kim, Byoungjae Min, Kennedy Edemacu, Yoonhyuk Choi, Sae-Hong Cho, Beakcheol Jang

首次发表
浏览论文内容

中文总结 AI 辅助

针对长期LLM智能体的自适应对话记录隐私威胁,提出差分隐私内存接口DP-MemView,经实验验证可在保留个性化和响应质量的同时降低隐私泄露风险。

中文摘要 AI 辅助

长期内存使大语言模型(LLM)智能体能够实现持续个性化,但即使未明确声明,基于内存的重复响应也可能累积泄露受保护属性。我们将此威胁形式化为自适应对话记录隐私,并引入DP-MemView,这是一种差分隐私接口,可私密选择用于响应条件化的公开视图,并将这些视图(而非原始内存)暴露给响应LLM。每次私密选择都会向其内存组与读取集相交的每个受保护属性收费。每个属性的账本会阻止任何超出其上限的选择,而是返回固定的通用视图。在明确的接口契约下,我们证明了整个自适应对话记录满足纯B_a-差分隐私(DP)。我们还将结果扩展到跨多个受保护组存在差异的存储,并限定观察对话记录可如何改变对手的先验概率。我们在受控相邻存储基准和公开语料迁移轨迹上,使用三个响应LLM评估了在线模式和预分配模式。两种模式均保持对话记录的可区分性接近随机水平,同时保留目标所需的个性化和整体响应质量。进一步诊断显示,移除关键安全措施会导致输出支持不匹配、缺少账本收费、泄露侧信道或长期泄漏增加。

英文摘要

Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly. We formalize this threat as adaptive transcript privacy and introduce DP-MemView, a differentially private interface that privately selects public response-conditioning views and exposes those views---rather than raw memory---to the response LLM. Each private selection is charged to every protected attribute whose memory group intersects the read set. Per-attribute ledgers block any selection that would exceed its cap and return a fixed generic view instead. Under an explicit interface contract, we prove pure B_a-DP for the entire adaptive transcript. We also extend the result to stores that differ across multiple protected groups and bound how much observing the transcript can change an adversary's prior odds. We evaluate the online and preallocated modes with three response LLMs on a controlled adjacent-store benchmark and a public-corpus transfer track. Both modes keep transcript distinguishability near chance while preserving target-required personalization and overall response quality. Further diagnostics show that removing key safeguards causes mismatched output support, missing ledger charges, revealing side channels, or growing long-horizon leakage.

发表机构

  • Sangmyung University(祥明大学)
  • College of Staten Island, The City University of New York(纽约城市大学斯塔滕岛学院)
  • Sookmyung Women’s University(淑明女子大学)
  • Hansung University(汉城大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑