arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于稀疏约束修正流的开放集视觉文本取证

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

Jiangling Zhang, Shuxuan Gao, Zeyu Chen, Yichao Liu, Yu Zhou

arXiv 2608.02258首次发表:更新:

AI 中文总结

针对生成式AI视觉文本篡改的开放集取证问题,提出基于SC-RF的生成式检测器,在三基准上F1、IoU优于亚军,零样本性能强,还可用于漏洞分析。

AI 中文摘要

快速发展的生成式AI能实现复杂的视觉文本篡改,且越来越难以被现有取证检测器检测到。现有的判别模型常过拟合特定的伪造模式,限制了其对未见过的开放集攻击的泛化能力。为解决这一挑战,我们提出一种生成式检测器,通过估计使查询图像与真实视觉文本统计量对齐所需的局部修复成本来定位篡改,而非学习伪造特定决策边界。具体而言,我们引入稀疏约束修正流(Sparse-Constraint Rectified Flow, SC-RF),这是针对空间稀疏异常定位的流匹配的检测器导向适配。我们还通过自监督的人工痕迹注入缓解数据稀缺问题,并使用像素空间取证DiT(Forensic-DiT)保留高频取证痕迹。在三个基准上的大量实验表明,我们的方法达到了最先进的性能,F1值和IoU分别比亚军高出3.2和4.8个百分点。特别是,所提出的检测器在具有挑战性的未见过的文本编辑模式上表现出强大的零样本性能。我们还提供了辅助压力测试分析,表明我们的模型产生的局部协调可以削弱现有检测器所依赖的统计线索,提供了互补的漏洞分析视角。

英文摘要

Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit specific forgery patterns, limiting their generalization to unseen, open-set attacks. To address this challenge, we propose a generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries. Specifically, we introduce Sparse-Constraint Rectified Flow (SC-RF), a detector-oriented adaptation of Flow Matching for spatially sparse anomaly localization. We further mitigate data scarcity via self-supervised Artifact Injection and preserve high-frequency forensic traces using a pixel-space Forensic-DiT. Extensive experiments on three benchmarks show that our method achieves state-of-the-art performance, surpassing the runner-up by 3.2 and 4.8 percentage points in F1 and IoU, respectively. In particular, the proposed detector demonstrates strong zero-shot performance on challenging unseen text editing patterns. We further provide an auxiliary stress-test analysis showing that local harmonization produced by our model can weaken the statistical cues relied upon by existing detectors, offering a complementary vulnerability-analysis perspective.

CommentsAccepted to ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑