arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13112cs.CV

面向物理保真的科学图表生成

Towards Physics-Faithful Generation of Scientific Diagrams

Minghui Zhang, Jinxin Shi, Yifan Chang, Liangliang Zhao, Yuandong Pu, Qian Yu, Ming Hu, Hanxiao Zhang, Yun Gu, Yirong Chen, Yu Qiao, Bo Zhang, Xiangchao Yan, Bin Fu, Yihao Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有文本到图像生成模型生成的科学图表存在物理错误的问题,提出Princigram生成器,通过结构化物理思维链实现物理保真的科学图表生成,在相关基准上验证了其有效性。

中文摘要 AI 辅助

文本到图像生成已达到照片级真实质量,但最先进的系统在生成科学图表方面仍不可靠,科学图表的价值不取决于外观,而取决于物理保真度:正确的力方向、有效的坐标系、一致的热力学状态以及与所描绘场景匹配的方程。通用模型在带有浅显物理描述的网络图像上训练,生成的图表看似合理,但存在物理错误,对教育和科学传播有害。我们提出了Princigram,一种物理保真的科学图表生成器及其数据流水线。我们的核心进展是结构化物理思维链(SP-CoT):一种针对各子学科的模式,将物理图表分解为跨六个子学科的显式多步推理链,从场景识别到力或过程分析,再到支配定律和合成。与自由形式的思维链不同,SP-CoT遵循带有严格保真规则的固定模式,将视觉基础事实与物理推断推理分开,并对所有数学进行符号化;它既作为密集训练监督,又在推理时作为结构化的“思考”提示。借助该方法,我们整理并结构化标注了430万张物理图像,其中115037张带有专家级标注,并适配了统一多模态主干。我们进一步推出VeriphyT2IBench,其问题源自每个预留图表自身的结构化标注:每个图表成为关于其对象、力和状态的特定于项目的二元问题库,因此评判模型的分数可分解为命名的物理事实,而非单一整体数字。在GenExam的物理子集和VeriphyT2IBench上,Princigram表明显式物理结构化监督可提高生成科学图表的物理保真度。

英文摘要

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

发表机构

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑