arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LeakageBench:文档图像中个人身份信息脱敏的文档级泄露风险基准

LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov

arXiv 2609.02207首次发表:更新:

发表机构

Center for Artificial Intelligence and Robotics (CAIRO); Technical University of Applied Sciences Würzburg-Schweinfurt (THWS); DataX(人工智能与机器人中心(开罗); 维尔茨堡-施韦因富特应用技术大学; DataX)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有PII基准未覆盖文档级脱敏风险的问题,推出LeakageBench基准,评估多种PII检测方法,发现虽工具辅助可提升定位效果,但多数页面仍存在高泄露风险。

AI 中文摘要

实际场景中的个人身份信息(PII)脱敏通常针对文档图像——扫描件、截图和PDF渲染图,其中光学字符识别(OCR)错误、布局结构和视觉噪声决定敏感信息是否被真正移除。现有PII基准大多以文本为中心,未衡量文档级脱敏风险:只要遗漏一个标识符,页面就仍不安全。我们推出LeakageBench,这一包含500张文档图像的挑战集,带有11954个符合GDPR的PII标注,涵盖直接标识符、关联键和上下文重识别表面。我们使用实体级F1、分组泄露和文档级泄露指标,评估通用OCR流水线、商业及任务适配的依赖OCR的检测器,以及无OCR的视觉语言模型。代码解释器将GPT-5.5的定位F1从0.090提升至0.249,但关键的页面级泄露仍达0.968。这些结果表明,更强的检测能力和工具辅助可提升定位效果,但无法让大多数页面安全发布。LeakageBench为文档图像中高召回、空间定位的PII脱敏提供诊断基准。

英文摘要

Real-world personally identifiable information (PII) redaction often operates on document images---scans, screenshots, and PDF renderings---where OCR errors, layout structure, and visual noise determine whether sensitive information is actually removed. Existing PII benchmarks are mostly text-centric and do not measure document-level redaction risk: a page remains unsafe if even one identifier is missed. We introduce LeakageBench, a challenge set of 500 document images with 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces. We evaluate generic OCR pipelines, commercial and task-adapted OCR-dependent detectors, and OCR-free vision-language models using entity-level F1, group-wise leakage, and document-level leakage metrics. Code Interpreter raises GPT-5.5 localization F1 from 0.090 to 0.249, but critical page-level leakage remains 0.968. These results show that stronger detection and tool assistance improve localization without making most pages safe for release. LeakageBench provides a diagnostic benchmark for high-recall, spatially grounded PII redaction in document images.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑