arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图像旁的文本:可信多模态医学数据中的检测、效用与泄漏及其更广泛领域

The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond

Andreas Maier, Monica Hinrichs-Mayer, Franziska Weber, Niklas Lackner, Matthias May, Bernhard Kainz, Siming Bayer

arXiv 2609.33280首次发表:更新:

AI 中文总结

本文衡量医学图像伴随文本的标识符检测、效用与身份泄漏,发现固定13检测器联合高敏感度,但跨文档链接仍能恢复身份,强调文本保护的必要性。

AI 中文摘要

医学图像与其描述性报告一同发布,保护图像并不能保护报告。本文衡量此类发布中的文本组成部分。我们在同一文档上测量标识符检测、下游效用和残余身份泄漏,以假名化策略作为测试变量:15个检测器、三种发布条件和四个语料库,涵盖医学报告、法律判决、新闻及其他体裁以及电子邮件,涉及德语、英语、中文和阿拉伯语。一个固定的13检测器联合在医学报告上达到0.9998的人员敏感度(特异性为0.8686),在法律判决上为0.9958(特异性0.8504),在新闻及其他体裁上为0.9352(特异性0.9318),在电子邮件上为0.9906(特异性0.6235)。使用该集成,与公开姓名列表的频率匹配在四个语料库中通过比对恢复的身份为零;其正确识别的姓名是检测器遗漏并保留在明文中的。跨文档链接在无训练情况下对0.93%的电子邮件查询将正确人员排在首位,有训练时为3.94%,而随机概率为1/3697,未修改文本上为71.98%。在医学报告上,无训练时未恢复任何身份,有训练时在138个查询中恢复0.71%,而随机概率为1/207,未修改文本上的上限为2.73%。

英文摘要

Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the variable under test: 15 detectors, three release conditions and four corpora of medical reports, legal judgments, news and other genres, and e-mail, in German, English, Chinese and Arabic. A fixed 13-detector union reaches a person sensitivity of 0.9998 at specificity 0.8686 on the medical reports, 0.9958 at 0.8504 on the legal judgments, 0.9352 at 0.9318 on news and other genres, and 0.9906 at 0.6235 on e-mail. With this ensemble, frequency matching with a public name list recovers zero identities by alignment across the four corpora; the names it got right were ones the detector missed, left in clear text. Cross-document linkage ranks the correct person first for 0.93% of e-mail queries without training and 3.94% with it, against 1/3697 chance and 71.98% on unmodified text. On the medical reports it recovers nothing without training and 0.71% of 138 queries with it, against 1/207 chance and a 2.73% ceiling on unmodified text.

Comments12 pages, submitted for peer review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑