发表机构
Faculty of Engineering, Shenzhen MSU-BIT University; Faculty of CMC, Shenzhen MSU-BIT University; Department of CDS, Indian Institute of Science(深圳北理莫斯科大学工程学院; 深圳北理莫斯科大学计算机、数学与力学学院; 印度科学学院计算机与数据科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对文档理解MLLMs的KIE任务,揭示其在视觉证据不足时会通过记忆的字段关系泄露隐私,提出DRUF框架与DocPrivacyBench基准,实验显示DRUF可提升泄露抑制效果并维持KIE性能。
AI 中文摘要
尽管多模态大语言模型(MLLMs)的隐私风险已引发广泛关注,但特定领域MLLMs的独特漏洞仍未得到充分探索。本文聚焦于用于身份证件处理的文档理解MLLMs,研究了关键信息提取(KIE)任务中固有的隐私问题。我们发现,当输入图像缺乏足够视觉证据时,这些模型往往依赖训练数据中记忆的字段关系来推断缺失内容,从而泄露包含敏感个人信息的多个关联字段。为缓解这一风险,我们提出三项关键贡献:其一,构建动态关系遗忘框架(DRUF),该框架包含关系解耦遗忘(RDU)模块与动态集合更新机制,可在抑制高风险字段对泄露的同时保留KIE性能;其二,引入DocPrivacyBench,这一新型基准可在视觉证据缺失或极少的条件下系统评估模型对隐私泄露的敏感性;其三,利用该基准评估三种MLLMs与六种遗忘方法,综合考量遗忘后的泄露抑制效果与实用性。实验结果表明,现有MLLMs在视觉证据匮乏时,尤其在噪声较大的数据集上,均会出现隐私泄露;相比之下,DRUF较最强基线方法的泄露抑制效果提升4.8个百分点,在维持稳健的文档信息提取性能的同时,有效缓解了隐私风险。
英文摘要
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer missing content, thereby leaking multiple correlated fields containing sensitive personal information. To mitigate this risk, we make three key contributions.First, we propose the Dynamic Relational Unlearning Framework (DRUF) which comprises a Relational Decoupling Unlearning (RDU) module and a dynamic set update mechanism. It suppresses the leakage of high-risk field pairs while preserving KIE performance.Second, we introduce DocPrivacyBench, a novel benchmark to systematically evaluate a model's susceptibility to privacy leakage under conditions of absent or minimal visual evidence.Third, we evaluate three MLLMs and six unlearning methods using this benchmark, assessing both post-unlearning leakage suppression and utility preservation.Our results demonstrate that existing MLLMs consistently exhibit privacy leakage when visual evidence is scarce, particularly on noisier datasets. In contrast, DRUF outperforms the strongest baseline by improving leakage suppression by 4.8 percentage points, effectively mitigating privacy risks while maintaining robust document information extraction performance.
CommentsACM mm 2026