AI 中文总结
本文提出SciFigPlag-Bench基准,用于科学图片的来源感知抄袭检测,含四类任务,实验发现视觉语言模型在细粒度来源推理等方面仍有挑战。
AI 中文摘要
科学图片往往编码着科学发现背后的视觉证据,但图片抄袭作为一个基准化的多模态评估问题仍未得到充分探索。本文提出SciFigPlag-Bench,这是一个针对学术文档中科学图片的来源感知推理基准。与通用图像相似度或图像取证基准不同,SciFigPlag-Bench评估可疑图片是否复用了特定来源图片的证据、复用内容如何被转换以及复用证据出现在何处。我们引入一种分解分类法,将复用内容与转换方式分开,涵盖材料保留型复用(如整图和子图复用)以及抽象内容型复用(如数据重表达和结构重绘)。基于该分类法,我们构建了一个混合基准,包含2582个正例对和2541个负例对,结合了已记录的真实案例、分类法引导的合成示例以及视觉相似的负例。该基准支持四项诊断任务:成对检测、来源归因、层级复用类型分类和复用对应定位。对多种视觉语言模型的实验建立了初始基线,并揭示了细粒度来源推理、复用类型理解和空间证据定位方面的持续挑战。
英文摘要
Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific source figure, how the reused content has been transformed, and where the reused evidence appears. We introduce a factorized taxonomy that separates what is reused from how it is transformed, covering material-preserving reuse, such as full-figure and subfigure reuse, as well as abstract-content reuse, such as data re-expression and structural redraw. Guided by this taxonomy, we construct a hybrid benchmark with 2,582 positive pairs and 2,541 negative pairs, combining documented real-world cases, taxonomy-guided synthetic examples, and visually similar negatives. The benchmark supports four diagnostic tasks: pairwise detection, source attribution, hierarchical reuse-type classification, and reuse correspondence localization. Experiments with diverse vision-language models establish initial baselines and reveal persistent challenges in fine-grained provenance reasoning, reuse-type understanding, and spatial evidence grounding.
Comments30 pages, 18 figures