arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种针对隐私保护大语言模型中嵌入到嵌入混淆的新型语义流形对齐攻击

A Novel Semantic Manifold Alignment Attack against Embedding-to-Embedding Obfuscation in Privacy-Preserving LLMs

Sicong Li, Lingfeng Yao, Xingke Yang, Ke Tu, Chenhao Wu, Hao Wang, Jiang Liu, Phone Lin, Xin Fu, Miao Pan

arXiv 2609.06749首次发表:更新:

发表机构

University of Houston; The Chinese University of Hong Kong, Shenzhen; Waseda University; Stevens Institute of Technology; National Taiwan University(休斯顿大学; 香港中文大学(深圳); 早稻田大学; 史蒂文斯理工学院; 国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对隐私保护LLMs中的嵌入到嵌入混淆,提出代理流形对齐攻击,利用语义结构保持特性,通过翻译任务重构明文,实验显示恢复性能优于现有方法。

AI 中文摘要

随着大语言模型(LLMs)的广泛应用,隐私保护推理对于敏感查询变得日益重要。为平衡隐私与效用,近期提出了一系列轻量级混淆方法,用户可在本地将明文嵌入转换为固定的密文嵌入。尽管此类嵌入到嵌入混淆(E2EO)方案在抵御传统词频攻击和嵌入反演攻击方面表现出相当的鲁棒性,但其核心机制仍是大规模的一对一替换,不提供任何密码学保证。本文提出代理流形对齐(PMA),一种针对隐私保护LLMs中E2EO的新型攻击。我们的关键观察是,E2EO方案保持了原始语义结构,因此混淆向量流可被视为一种未知分词器语言,其符号即为向量本身。据此,所提出的密文到明文重构攻击可表述为从该未知分词器语言到明文的翻译任务。具体而言,仅通过访问混淆向量流、目标分词器及公开语料库,PMA攻击首先利用Word2Vec分别对混淆流和公开语料库中的共现模式进行建模,构建两个代理向量嵌入。随后,攻击基于结构相似性对齐这两个嵌入的底层流形。最后,将混淆向量映射回明文。实验结果表明,PMA在明文恢复方面始终优于其他最先进的攻击方法。

英文摘要

With the widespread applications of large language models (LLMs), privacy-preserving inference has become increasingly essential for sensitive queries. To balance privacy and utility, a series of lightweight obfuscation approaches has recently been proposed, where users locally transform plaintext embeddings into the fixed ciphertext ones. While such Embedding-to-Embedding Obfuscation (E2EO) schemes demonstrate considerable resilience against traditional token frequency and embedding inversion attacks, the core mechanism behind remains to be the large-scale one-to-one substitution, which provides no cryptographic guarantees. In this paper, we propose Proxy Manifold Alignment (PMA), a novel attack against E2EO in privacy-preserving LLMs. Our key observation is that E2EO schemes keep the original semantic structure, so that the obfuscated vector stream can be regarded as an unknown tokenizer-language whose symbols are the vectors themselves. Therefore, the proposed ciphertext to plaintext reconstruction attack can be formulated as a translation task from the unknown tokenizer-language to plaintext. Specifically, by only accessing the obfuscated vector stream, the target tokenizer and a public corpus, the PMA attack first employs Word2Vec to model the co-occurrence patterns within the obfuscated stream and the public corpus independently, and constructs two proxy vector embeddings. Then, the attack aligns the underlying manifolds of these two embeddings based on structural similarity. Finally, it maps the obfuscated vectors back to plaintext. Experimental results demonstrate that PMA consistently achieves higher plaintext recovery than other state-of-the-art attack methods.

CommentsAccepted at the EMNLP 2026 Main Conference as a long paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑