arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeepInvert:针对混淆语言模型的半监督嵌入反演攻击

DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

Zhicong Huang, Cheng Hong, Tao Wei

arXiv 2608.04477首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对ObfusLM等混淆语言模型防御提出半监督嵌入反演攻击DeepInvert,其标记恢复率达73.5%,优于现有方法,揭示混淆防御存在效用与反演的权衡问题,呼吁重新评估该防御类别。

AI 中文摘要

基于云的语言模型服务通常会处理包含敏感信息的提示。基于混淆的防御措施——包括ObfusLM、SentinelLMs、TextObfuscator和DPNR——通过在传输前转换提示表示来降低这种风险,为密码学解决方案提供了一种轻量级替代方案。我们表明,这些防御措施提供的保护远低于此前认为的水平。我们提出了DeepInvert,这是一种半监督嵌入反演攻击,可从混淆表示中恢复原始标记,其精度高于现有方法。关键见解是,尽管存在扰动,未标记的混淆嵌入仍保留可利用的语义结构。DeepInvert结合了对标记影子数据的监督训练,以及对未标记目标嵌入的新型无监督一致性目标,通过混合训练流程在两者之间交替。针对防御的适配进一步将攻击扩展到基于编码器和自回归架构的多种混淆机制。在9种防御、5个任务和4种模型架构上的实验表明,DeepInvert在大多数防御上的性能优于现有攻击。针对ObfusLM,DeepInvert实现了73.5%的Top-1标记恢复率,而之前的最佳方法仅为26.2%。我们的结果揭示了一种依赖于任务的权衡:保留足够效用信号的混淆方案也保留了足够的结构以进行反演,而抵抗反演的方案则会破坏效用。在更简单的分类任务上,一些基于差分隐私(DP)的防御可以同时维持两者。我们呼吁对这一防御类别进行重新评估。

英文摘要

Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, SentinelLMs, TextObfuscator, and DPNR---mitigate this risk by transforming prompt representations before transmission, offering a lightweight alternative to cryptographic solutions. We show these defenses provide far less protection than previously believed. We present DeepInvert, a semi-supervised embedding inversion attack that recovers original tokens from obfuscated representations with higher accuracy than prior methods. The key insight is that unlabeled obfuscated embeddings retain exploitable semantic structure despite perturbation. DeepInvert combines supervised training on labeled shadow data with a novel unsupervised consistency objective over unlabeled target embeddings, alternating between the two via a mixed training pipeline. Defense-aware adaptations further extend the attack to diverse obfuscation mechanisms across encoder-based and autoregressive architectures. Experiments on nine defenses, five tasks, and four model architectures show that DeepInvert outperforms prior attacks on most defenses. Against ObfusLM, DeepInvert achieves 73.5\% top-1 token recovery versus 26.2\% for the previous best. Our results reveal a task-dependent tension: obfuscation schemes preserving enough signal for utility also retain sufficient structure for inversion, while schemes resisting inversion collapse utility. On simpler classification tasks, some DP-based defenses can maintain both. We call for a re-evaluation of this defense class.

Comments20 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑