发表机构
University of Rome Tor Vergata; University of Luxembourg; Almawave Labs(罗马第二大学; 卢森堡大学; Almawave 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明训练数据提取攻击可跨语言迁移,多语言模型在非英语提示下仍能泄露记忆的PII,且中间层存在跨语言桥梁,凸显需语言无关的净化策略。
AI 中文摘要
大型语言模型(LLMs)中个人身份信息(PII)保护的鲁棒性是一个关键问题,然而与跨语言数据提取相关的风险仍未得到充分探索。本研究评估了以英语为中心和多语言的模型在非英语语言提示下对训练数据提取(TDE)攻击的脆弱性。我们构建了一个包含社交媒体账号、电子邮件地址和电话号码的多领域PII数据集,并将攻击上下文翻译成意大利语、西班牙语、法语和德语。我们的结果表明,针对以英语为中心和多语言模型的TDE攻击可迁移到不同语言:攻击在翻译后的提示上成功,即使只有原始英语提示可能包含在预训练数据中。对部分翻译样本的网页存在性检查确认它们无法在线获取。在其他语言中恢复的英语泄露比例随模型的多语言能力增强而增加,而当原始措辞丢失时,即使语言不变,该比例也急剧下降。这表明原生多语言预训练促进了潜在跨语言桥梁的出现,从而简化了个人身份信息(PII)的检索。我们分析了多语言大型语言模型(LLMs)的激活,发现同一提示的不同翻译在相似表示中被桥接,其中中间层的对齐最强。我们的结果突显了现代LLMs中的根本性安全漏洞,需要更鲁棒的、与语言无关的净化策略用于未来模型对齐。
英文摘要
The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.