发表机构
Center for Applied Intelligent Systems Research, Halmstad University(哈尔姆斯塔德大学应用智能系统研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示三类主流PII检测系统在分布外输入下存在显著鲁棒性缺陷,提出带问答驱动反馈循环的混合检测流水线,并发布相关压力测试基准。
AI 中文摘要
个人身份信息(PII)检测是数据保护基础设施的核心组成部分,漏检实体将直接构成隐私与安全风险。尽管现代PII系统在标准基准测试中表现出强劲性能,但我们发现这类评估掩盖了其在部署时遇到的真实分布偏移下的严重鲁棒性缺陷。我们未比较各模型的最先进准确率,而是研究不同PII检测范式在噪声、非结构化及非正式输入下的失效情况。我们构建了涵盖七类自然分布偏移的压力测试基准,并评估了三类广泛部署的架构家族的代表性系统:基于编码器的命名实体识别(SpaCy)、基于规则的混合检测(Presidio)以及生成式大语言模型提取(Qwen2.5-3B)。这三类系统在分布外输入上均出现显著性能下降,但失效模式各有不同且互为补充:编码器模型主要在未见表面形式与边界检测上失效,基于规则的系统在非标准格式上失效,大语言模型则表现出实体类型混淆与生成不稳定问题。这些结果表明,聚合基准分数掩盖了部署时的关键弱点,且没有单一架构在所有PII类别中都具备一致可靠性。基于上述发现,我们提出了一种带有问答驱动反馈循环的混合检测流水线,用于迭代缓解风险,并发布我们的基准以支持面向分布外的PII系统评估。
英文摘要
Personally Identifiable Information (PII) detection is a foundational component of data protection infrastructure where missed entities constitute direct privacy and security risks. Although modern PII systems report strong performance on standard benchmarks, we show that these evaluations mask substantial robustness failures under realistic distribution shifts encountered in deployment. Rather than comparing state-of-the-art accuracy, we study how different PII detection paradigms fail under noisy, unstructured, and informal inputs. We construct a stress test benchmark spanning seven categories of natural distribution shift and evaluate representative systems from three widely deployed architectural families: encoder-based NER (SpaCy), rule-based hybrid detection (Presidio), and generative LLM extraction (Qwen2.5-3B). All three exhibit significant degradation on out-of-distribution inputs, but with distinct and complementary failure modes. Encoder models primarily fail on unseen surface forms and boundary detection, rule-based systems fail on non-standard formats, and LLMs exhibit entity-type confusion and generation instability. These results show that aggregate benchmark scores obscure deployment-critical weaknesses and that no single architecture is uniformly reliable across PII categories. Motivated by these findings, we propose a hybrid detection pipeline with a QA-driven feedback loop for iterative risk mitigation, and release our benchmark to support OOD-aware evaluation of PII systems.