AI 中文总结
研究针对大语言模型生成技术内容常违反科学原理的问题,提出图结构共形预测框架SFC,将科学推理分解为单元,考虑逻辑依赖,通过实时验证和动态分支纠错,在多基准测试中表现出色,提升准确率、有效性并减少定律违反。
AI 中文摘要
大型语言模型在生成技术内容时经常违反基本科学原理,损害其在科学应用中的可靠性。我们引入了科学可行性控制(SFC),这是一个图结构的共形预测框架,通过渐进的绝对连贯事实性验证为科学推理有效性提供统计保证。该方法将科学推理分解为原子绝对连贯事实性单元,解决早期科学错误污染后续推理步骤的级联效应。与孤立处理断言的基于独立性的方法不同,SFC将逻辑依赖视为近似可推导性图,检测到科学违规时通过动态分支进行实时验证并转向替代生成路径。我们在多个科学推理基准上展示了SFC,在PhyX物理推理上达到50.1%的准确率,优于近期模型,同时提供91.7%的科学有效性并减少73%的科学定律违反。
英文摘要
Large language models frequently violate fundamental scientific principles when generating technical content, undermining their reliability in scientific applications. We introduce Scientific Feasibility Control SFC, a graph-structured conformal prediction framework that provides statistical guarantees for scientific reasoning validity through progressive absolute-coherent-factuality validation. Our approach decomposes scientific reasoning into atomic absolute-coherent-factuality units requiring both individual correctness against physical laws and logical substantiation from preceding context, addressing the cascade effect where early scientific errors contaminate subsequent reasoning steps. Unlike independence-based methods that treat claims in isolation, SFC models logical dependencies as approximate deducibility graphs and operates through real-time validation with dynamic branching when scientific violations are detected, the system branches to alternative generation paths using verified context as foundation. We demonstrate SFC across established scientific reasoning benchmarks including PhyX multimodal physics, MATH, ScienceQA, and ARC Challenge, achieving 50.1 percent accuracy on PhyX physics reasoning, substantially outperforming recent reasoning models including DeepSeek-R1 49.8 percent and GPT-4 45.8 percent while providing 91.7 percent scientific validity with formal conformal coverage guarantees at alpha equals 0.10 confidence level and reducing scientific law violations by 73 percent across multiple model architectures.
Comments25 pages