arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02616cs.CLcs.AI

评估OpenAI的隐私过滤器:在42个基准测试中开展跨语言、跨领域的个人身份信息(PII)检测

OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks

Rohith Uppala

首次发表
浏览论文内容

中文总结 AI 辅助

本研究首次系统评估OpenAI的15亿参数PII检测器OPF,在42个跨语言跨领域基准中对比其与Presidio、XLM-RoBERTa、GPT-4o的PII检测性能,揭示其优势与缺陷。

中文摘要 AI 辅助

我们首次对OpenAI的隐私过滤器(OpenAI's Privacy Filter,OPF)开展独立、系统的评估,该模型是一个拥有15亿参数的双向个人身份信息(PII)检测器,评估覆盖了22种语言、5个领域的42个人工合成基准测试。在零样本设置下,OPF在AI4Privacy基准测试上的F1值为0.855,在SPY医疗基准测试上的F1值为0.464,在带有PII标注的基准测试中,其表现优于Presidio(对应F1值分别为0.431、0.273)和XLM-RoBERTa(对应F1值分别为0.269、0.111);在多语言命名实体识别(NER)任务中,XLM-RoBERTa在所有13种印度语言及非拉丁语言上的表现均优于OPF。GPT-4o在医疗、法律和金融PII检测任务中表现领先(SPY基准测试平均F1值为0.643,Gretel基准测试平均F1值为0.527),而OPF在结构化合成PII检测(平均F1值为0.71)和客户支持场景(F1值为0.60)中表现领先。当PII嵌入叙事性散文中时,OPF的性能会急剧下降:在NER基准测试上的F1值为0.04至0.57,且在非拉丁文字(阿拉伯语:0.04,西里尔语:0.03)上完全失效。错误分析显示,OPF在结构规则的PII类型上表现最强(邮箱:0.78,电话:0.76),在文化可变的PII类型上表现最弱(人名:0.40,地址:0.49),且在客户支持及医疗/法律PII检测中存在召回率偏向(精确率P为0.31至0.54,召回率R为0.70至0.85);在所有领域的全局精确率范围为0.31至0.86。

英文摘要

We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that converts an autoregressive language model into a bidirectional PII detector, across 32 benchmarks spanning 14 languages and 5 domains. Our most practically actionable finding is a domain-dependent labeled-data crossover: fine-tuned XLM-RoBERTa surpasses OPF's zero-shot performance with only ~500 labeled examples on English synthetic PII (~100 on non-English Kiji), and ~1000 on synthetic medical PII. Crucially, per-class fine-tuning (17 PII entity types, a subset of OPF's 33) is less data-efficient than binary labels at small n -- at n=100, binary F1=0.634 vs. per-class 0.360. Zero-shot, OPF achieves F1=0.464 on the SPY medical benchmark and F1=0.855 on AI4Privacy, substantially outperforming Presidio and XLM-RoBERTa-large-NER. However, OPF degrades sharply outside its PII training distribution: F1=0.04--0.40 on general NER benchmarks and collapses for non-Latin scripts (Arabic: 0.04, Cyrillic: 0.03). Error analysis reveals OPF excels on structurally regular PII (email: 0.78, phone: 0.76) but struggles with culturally variable entities (person names: 0.40, addresses: 0.49), and is recall-biased across most PII domains (precision 0.31--0.54, recall 0.70--0.85). We provide a decision heuristic for when to use OPF zero-shot, when to fine-tune XLM-RoBERTa, and which language families to avoid.

补充信息

↑