MemLeak:多租户AI智能体记忆中的跨用户语义泄露
MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory
浏览论文内容
中文总结 AI 辅助
针对多租户AI智能体共享向量存储导致的跨用户语义泄露问题,提出形式化定义并实验验证泄露严重性,发现硬性检索后所有权门控能有效缓解,恢复干净基线。
中文摘要 AI 辅助
在企业多租户部署中,个人AI智能体共享一个用于长期记忆的向量存储库。共享的嵌入空间为跨用户记忆泄露创造了表面:用户的查询可以通过普通的余弦相似度检索,在没有任何利用手段的情况下,检索到属于另一用户的语义相邻记忆。我们将此形式化为跨用户可接受性失败,并在六项实验及后续消融实验中,分别在稀疏(TF-IDF)和生产级(MiniLM-L6-v2)检索下进行评估。非对抗性的、偶然的泄露在池化{同团队}检索下达到70--100%;对抗性精心构造的记忆实现90--100%的top-$k$放置,超过了较弱的基于关键词的攻击者基线,在生产级密集检索(配置B)下得分提升为$+0.416$到$+0.511$;端到端响应污染在生产检索路径下达到5.00/5,在Claude Sonnet 4.5下达到4.67/5,且受污染的响应通常被评为与干净响应一样有帮助或更有帮助,这一差距已通过人类判断验证。在三种架构缓解措施中,只有硬性的检索后所有权门控能够一致地恢复干净基线(1.00/5),跨越{两个生成模型},且测量到的延迟开销约为每次查询1.4毫秒。
英文摘要
Personal AI agents in enterprise multi-tenant deployments share a common vector store for long-term memory. Shared embedding spaces create a surface for cross-user memory leakage: a user's query can retrieve semantically adjacent memories belonging to another user through ordinary cosine-similarity retrieval, without any exploit. We formalize this as cross-user admissibility failure and evaluate it across six experiments, plus follow-up ablations, under both sparse (TF-IDF) and production-faithful (MiniLM-L6-v2) retrieval. Non-adversarial, incidental leakage reaches 70--100\% under pooled {same-team} retrieval; adversarially crafted memories achieve 90--100\% top-$k$ placement, exceeding weaker keyword-based attacker baselines, with score lifts of $+0.416$ to $+0.511$ under production-faithful dense retrieval (Config B); and end-to-end response contamination reaches 5.00/5 under a production retrieval path and 4.67/5 with Claude Sonnet~4.5, with contaminated responses often scoring as helpful or more helpful than clean ones, a gap validated against human judgment. Among three architectural mitigations, only hard post-retrieval ownership gating consistently restores the clean baseline (1.00/5) across {two generation models, at a measured latency overhead of roughly 1.4~ms per query.
发表机构
- Workday AI Research(Workday AI 研究院)
机构由 AI 辅助整理,请以论文原文为准。