发表机构
Shenzhen University; Tencent Youtu Lab(深圳大学; 腾讯优图实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对篡改文本检测中专家模型泛化差与MLLM不敏感的问题,提出专家知识内化(EKI)两阶段框架,通过空间聚焦与表示对齐将取证能力内化,实现SOTA性能且推理无需外部专家。
AI 中文摘要
篡改文本检测(TTD)对于在安全关键工作流程中保障文档真实性至关重要。现有的专家模型能够有效捕捉细微的篡改痕迹,但在不同文档领域间往往泛化能力较差,而多模态大语言模型(MLLMs)则提供了更强的语义理解和迁移能力,但对细粒度的取证痕迹仍不敏感。这种互补性促使我们研究如何将专家取证感知内化到MLLM中,而不仅仅是通过外部模块访问。我们识别出阻碍这一目标的一个根本性“双重不匹配”:一是粗粒度视觉标记与微小篡改区域之间的空间精度不匹配,二是面向语义的预训练与低层级取证感知之间的感知粒度不匹配。为解决这些挑战,我们提出了专家知识内化(EKI),这是一个渐进式两阶段框架,将取证专业知识迁移到MLLM本身。在第一阶段,文本聚焦和图像聚焦策略建立了对小型文本区域的精确空间聚焦。在第二阶段,所提出的取证通用表示对齐(FGRA)损失将浅层LLM表示与预训练取证专家的表示对齐,使模型能够在这些线索被更深层的语义抽象稀释之前获得细粒度的伪影感知。在多个域内和跨域基准上的大量实验表明,EKI达到了最先进的性能,并且比现有的基于专家模型和基于MLLM的方法具有更强的泛化能力。此外,专家仅在训练时需要,使得最终的MLLM在推理时无需依赖任何外部专家,即可保持与原始模型几乎相同的推理效率。
英文摘要
Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.