英语源多语言检索增强生成(RAG)系统的隐私风险所在:五种查询语言下的阶段分解审计
Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages
浏览论文内容
中文总结 AI 辅助
该研究针对英语源多语言RAG系统,通过含五种查询语言的两阶段防御审计发现,非英语语言仍存在残留PII泄露,附加黄金文档可阻断多数残留单元,后续计划开展独立MT与非Qwen判别器的复现工作。
中文摘要 AI 辅助
普遍假设认为,切换到非英语语言会使多语言检索增强生成(RAG)系统更容易遭受个人信息攻击。我们在英语源合成个人身份信息(PII)语料库上对这一假设进行了测试,该语料库包含五种查询语言,且采用两阶段防御机制(大语言模型(LLM)输入判别器 + 正则表达式输出过滤器),其翻译器、判别器、回译器和生成器均为Qwen2.5-7B——因此以下所有发现均是该 pipeline 条件下的结果,而非语言固有风险的因果排名。在仅输出过滤的情况下,英语的非结构化PII泄露率最高;在文档级自助抽样区间下,仅英语与斯瓦希里语存在显著差异。一旦加入输入判别器,阿拉伯语和斯瓦希里语仍存在残留泄露,且回译查询并不能缩小差距(我们报告了这一消融实验结果,但不能将其用作因果诊断,因为回译器同样是Qwen模型)。在一个独立的n=17多语言提示判别器残留角落案例中,将黄金语料库文档附加到输入判别器可阻断15/17个残留单元。我们将这一结果视为机制诊断,而非可部署的防御措施:它使用 oracle 检索,阻断/允许率仅针对对抗性查询测量,且我们未测量良性查询的误报率和答案效用成本。补充材料包含代码、语料库、查询及每次试验的JSONL;优先后续工作是使用独立机器翻译(MT)加非Qwen判别器,结合母语者查询集进行复现,其范围在局限性部分明确。
英文摘要
A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test this on an English-source synthetic-PII corpus with five query languages and a two-stage defence (LLM input judge + regex output filter), in a pipeline whose translator, judge, back-translator, and generator are all Qwen2.5-7B -- so every finding below is pipeline-conditional, not a causal ranking of language-inherent risk. Under output-only filtering, English has the highest observed unstructured-PII leak rate; only English-vs-Swahili separates cleanly under document-level bootstrap intervals. Once the input judge is added, residual leaks remain on Arabic and Swahili, and back-translating the query does not close the gap (an ablation we report but cannot use as a causal diagnostic, since the back-translator is also Qwen). On a separate n=17 multilingual-prompted-judge residual corner, attaching the gold corpus document to the input judge blocks 15/17 residual cells. We frame this last result as a mechanism diagnostic, not a deployable defence: it uses oracle retrieval, BLOCK/ALLOW rates are measured on adversarial queries only, and we measure no benign-query false-positive rate and no answer-utility cost. The supplementary material contains code, corpora, queries, and per-trial JSONLs; the priority follow-up is an independent-MT plus non-Qwen-judge replication with a native-speaker query set, scoped in the Limitations section.
发表机构
- Northeastern University(东北大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Southern Methodist University(南卫理公会大学)
机构由 AI 辅助整理,请以论文原文为准。