AI 中文总结
本文构建了含3857张科学图表的标注数据集,提出基于稿件证据的多模态评分器SciFigAlign,其在测试集上的MAE为0.3524,相对最佳LLM基准降低59%误差,证实需视觉与稿件证据的学习对齐。
AI 中文摘要
同行评审中的科学图表评估与通用图像质量评估存在根本差异:图表必须视觉清晰、忠实支持稿件主张,并通过清晰的视觉层级传递证据。然而,将传统图像评估方法应用于科学图表质量评估时会暴露出局限性:经典图像质量评估(IQA)模型仅能捕捉感知质量或美学,无法判断图表是否服务于论文的科学论证;基于CLIP的方法可评估通用图像-文本对应关系,但缺乏对稿件上下文的理解;零样本大语言模型(LLM)/视觉语言模型(VLM)裁判被用于图表评分时,往往会产生过于集中的分数,且视觉与文本证据的融合有限。本文构建了一个包含3857张来自同行评审会议论文的科学图表的标注数据集,每张图表均按同行评审导向的四个维度评分:清晰度、相关性、信息量和结构。我们提出SciFigAlign,这是一种基于稿件证据的微调多模态评分器,用于评估图表质量。给定图表裁剪图、标题、引用段落及精简的论文上下文,SciFigAlign通过各模态交叉注意力机制和CubeMLP融合,端到端微调CLIP与SciBERT,联合优化SmoothL1损失与论文内排序铰链损失。在论文级划分的测试集(n=396)上,SciFigAlign的平均绝对误差(MAE)为0.3524,论文内成对准确率达81.64%,相比最佳LLM-as-judge基准(MAE=0.864)实现了59%的相对误差降低。消融实验证实,基于稿件的输入、引用上下文去噪及排序监督均至关重要,表明科学图表评估需要视觉内容与稿件证据间的学习对齐,而非仅通过提示,即便使用最先进的VLM亦是如此。
英文摘要
Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript context; and zero-shot LLM/VLM judges, when repurposed for figure scoring, often yield overly concentrated scores with limited fusion of visual and textual evidence. We introduce an annotated dataset of 3,857 scientific figures from peer-reviewed conference papers, each rated along four peer-review-oriented dimensions: Clarity, Relevance, Informativeness, and Structure. We propose SciFigAlign, a fine-tuned multimodal scorer that grounds figure quality assessment in manuscript evidence. Given a figure crop, caption, citing paragraphs, and light paper context, SciFigAlign fine-tunes CLIP and SciBERT end-to-end with per-modality cross-attention and CubeMLP fusion, jointly optimizing SmoothL1 regression with a within-paper ranking hinge loss. Under paper-level splits, SciFigAlign achieves a macro MAE of 0.3524 and a within-paper pairwise accuracy of 81.64% on the test set with n = 396, a 59% relative error reduction over the best LLM-as-judge baseline with MAE 0.864. Ablations confirm that manuscript-grounded inputs, citing-context denoising, and ranking supervision are all critical, showing that scientific figure assessment requires learned alignment between visual content and manuscript evidence rather than prompting alone, even with state-of-the-art VLMs.