arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciFigQual-Bench:面向全文本上下文的科学图片质量评估基准

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong, Weidong Cai, Lequan Yu

arXiv 2607.27084首次发表:更新:

AI 中文总结

本文针对现有科学图片质量评估方法的不足,构建了SciFigQual-Bench基准,设计SFQ-Agent框架实现自动化评估,实验显示其在eval1200子集上表现优于主流方案。

AI 中文摘要

科学图片是呈现实验结论、阐述系统架构及支撑科学论文中对比论证的核心要素。然而,现有图像质量评估(IQA)方法主要针对自然照片或AI生成内容设计,无法直接应用于科学论文。少数针对学术图表的研究仍局限于视觉表面对比,未能验证标题对齐度、引用相关性或视觉误导性。为解决该问题,本文提出SciFigQual-Bench,这是一个基于全文本上下文的基准,从清晰度、布局、标题适配度、上下文相关性和误导风险五个维度评估科学图片。数据涵盖2020至2025年计算机科学顶会,共6308张图片由多名领域专家对五个维度独立评分并聚合为金标准标注。与以往科学图片基准不同,该数据集将每张图片与其标题、引用句及论文上下文绑定。为实现该基准的自动化评估,本文设计了分阶段跨模态评估框架SFQ-Agent,通过收集和融合模态证据实现可审计的精细化评分。在测试子集eval1200上对多个主流大模型进行评估,配备GPT-5.6-Sol的SFQ-Agent (F3)取得最低总体平均绝对误差(0.418)和最高一致性率(93.4%),始终优于直接评估及辅助(Sidecar)视觉语言模型评估方案。

英文摘要

Scientific figure quality is bound to the manuscript: a crop can be visually clear and still contradict its caption or the paragraph that cites it. Most image quality assessment and chart-understanding methods target a detached visual surface or a question-answering score, rather than alignment among the figure, its caption, and the citing text. A single end-to-end prompt that sees the image and the text together still mixes what is visible with what the caption and the citing paragraph claim, so fluent wording can raise a visual score and missing text can be treated as low quality. To address this gap, we introduce SciFigQual-Bench, a manuscript-linked benchmark for published CS-conference figures that binds each figure to its caption and to index-resolved citing paragraphs, and scores visual clarity, layout, caption consistency, context consistency, and misleading risk, leaving a dimension unevaluated when its evidence is absent. SFQ-Agent reads the image and the text in separate calls and records modality-specific evidence. A cross-modal judge scores caption and citing-text alignment from those reports, and a deterministic runner keeps the visual scores, sets misleading risk by a fixed rule, and applies the written caps, so each dimension follows the rubric instead of a single end-to-end prompt. Experiments comparing this staged judge with single-pass and OCR-sidecar protocols find the closest point-estimate fit to mean human ratings under staged judging, while the gap between protocols remains small. Caption consistency remains the main gap, and agreement with the rater mean answers a different question from agreement among human raters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑