arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciGen-Verifier:科学图像生成中可解释验证的多模态推理器

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin, Xi Yu, Xusen Hei, DingBa Fu, Jiayuan Xie, Yi Cai

arXiv 2609.33399首次发表:更新:

发表机构

South China University of Technology; The Hong Kong Polytechnic University(华南理工大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SciGen-Verifier多模态推理验证器及SciGen-Verify基准,通过两阶段强化学习实现科学图像生成的可解释验证与纠错,性能媲美大型专有模型。

AI 中文摘要

在现实教育中,解决方案往往不仅用文字表达,还通过绘图表达——如电路图、几何构造、函数图像——而教师必须像批改文字一样仔细地批改这些绘图。近年来,统一多模态模型的进展使得科学图像生成成为可能,然而验证这些专业视觉输出的正确性仍然是一个关键瓶颈:错误往往源于复杂的领域知识、结构推理和多步骤指令,而非表面层面的伪影。现有的验证器主要针对自然图像,并将判断压缩为标量分数,导致科学覆盖和用于纠错的可解释反馈未被充分探索。为弥补这一差距,我们做出了三项主要贡献。(1)我们构建了SciGen-Verify基准,专门用于科学图像生成的可解释验证,涵盖指令遵循、多学科推理和世界知识领域。它包含一个三层级协议,基于二元判断,支持解释和纠正性编辑指令。(2)我们开发了SciGen-Verifier,一种推理驱动的多模态验证器,通过冷启动监督微调,随后进行基于课程的两阶段强化学习流水线进行训练。规则引导的过程首先通过奖励加强科学推理探索,随后通过结果奖励使输出与真实标注对齐。(3)在SciGen-Verify上,SciGen-Verifier取得了与更大的专有模型相比有竞争力的性能。它进一步作为实用的在线批评者,用于迭代图像修正。

英文摘要

In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑