arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38391cs.CV

MSU GenText-Forensics 挑战赛 2026 技术报告

Team MSU GenText-Forensics Challenge 2026 Technical Report

Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova

首次发表
浏览论文内容

中文总结 AI 辅助

针对文档伪造挑战,我们提出分解式思维链流水线,结合篡改检测器与两个 LoRA 适配的 Qwen3-VL-32B 模型,通过教师蒸馏训练,实现定位、分类与报告生成,获挑战赛第三名。

中文摘要 AI 辅助

文档文本伪造已超越简单的像素级操作:现代攻击不仅改变文档的外观,还改变其含义,并日益针对消费此类文档的 OCR 与 LLM 流水线。因此,ACM MM 2026 GenText-Forensics 挑战赛要求系统不仅判断多语言文本图像是否被伪造,还要定位篡改点、识别攻击类型,并生成带有支持证据的可读取证报告。我们提出了一种分解式思维链(CoT)流水线,结合文档篡改检测器(DTD)与两个 Qwen3-VL-32B 视觉语言模型,每个模型通过 LoRA 适配到不同的子任务。DTD 生成篡改概率图,并转换为编号的候选区域;第一个模型(过滤器)验证这些区域并分配初步的伪造类型,而第二个模型(语义侦探)合并并重新定位存活的区域,搜索像素级检测器不可见的纯语义异常,并撰写最终报告。两个模型均通过从特权 Qwen3-VL-235B 教师模型(可访问真实掩码和报告)中蒸馏思维链轨迹进行训练。我们的方法在 ACM MM 2026 GenText-Forensics 挑战赛中获得第三名。我们详细描述了数据准备、测试时增强、区域渲染、蒸馏协议和训练配置,并报告了检测器阈值、提示设计和流水线分解的消融实验。

英文摘要

Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR & LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.

发表机构

  • Lomonosov Moscow State University(莫斯科国立大学)
  • MSU Institute for Artificial Intelligence(莫斯科国立大学人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑