arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciFigure2Code:一个用于科学图表到代码的AI重建基准

SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code

Wentao Li, Yibo Wu, Yizhe Chen, Ruixuan Chen, Jiangjie Qiu, Yijun Li, Zhao Leyi, Xiaonan Wang

arXiv 2609.08155首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SciFigure2Code基准,通过AI重建和审计协议,将科学图版转化为可编辑Python程序,并构建包含6,740个图版和337图版测试集,评估14个零样本模型,发现执行和布局仍是弱点。

AI 中文摘要

科学图表是研究声明被检查和复用的界面,但最终发表的图版很少公开生成它们的数据或绘图代码。因此,从像素中恢复这种隐藏的出处是欠定问题。我们引入了SciFigure2Code,一个AI重建的基准,它转而评估呈现恢复:生成可编辑的Python程序,保留科学图版的排列和阅读方式。角色专用的Codex智能体生成、执行、视觉优化并审计银标准呈现程序,这些程序捕获几何、视觉层次、编码、注释和排版,而不声称恢复原始测量数据或作者源代码。这种重建与审计协议将最终发表的图版转化为可审计的参考包;由此产生的资源包含6,740个经审阅的图版和SciFigureBench,一个平衡的337图版测试集,涵盖31种图表子类型、五个领域和三个复杂度级别。在仅图像和带标题辅助设置下的14个零样本模型中,执行、多组件布局、坐标轴、图例和科学标签仍然薄弱。Claude Opus 4.7在仅图像总体得分中最高,Claude Opus 4.6在带标题辅助重建中领先,而两阶段先规划后编码的提示方法提高了所有四个测试模型的总体得分。SciFigure2Code为构建可编辑、视觉忠实的科学图表呈现的智能体提供了一个可审计的测试平台。

英文摘要

Scientific figures are the interface through which research claims are inspected and reused, but final published panels rarely expose the data or plotting code that produced them. Recovering this hidden provenance from pixels is therefore underdetermined. We introduce SciFigure2Code, an AI-reconstructed benchmark that instead evaluates presentation recovery: generating editable Python programs that preserve how a scientific panel is arranged and read. Role-specialized Codex agents generate, execute, visually refine, and audit silver-standard presentation programs that capture geometry, visual hierarchy, encodings, annotations, and typography without claiming to recover original measurements or author source code. This reconstruction-and-audit protocol turns final published panels into auditable reference packages; the resulting resource contains 6,740 reviewed panels and SciFigureBench, a balanced 337-panel test set across 31 chart subtypes, five domains, and three complexity levels. Across 14 zero-shot models in image-only and caption-assisted settings, execution, multi-component layouts, axes, legends, and scientific labels remain weak. Claude Opus 4.7 achieves the highest image-only Overall score, Claude Opus 4.6 leads caption-assisted reconstruction, and two-stage plan-then-code prompting improves Overall for all four tested models. SciFigure2Code provides an auditable testbed for agents that construct editable, visually faithful scientific figure presentations.

Comments20 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑