发表机构
University of Pennsylvania; William & Mary; University of Illinois Urbana-Champaign; Amazon; Stevens Institute of Technology; Northwestern University(宾夕法尼亚大学; 威廉玛丽学院; 伊利诺伊大学厄巴纳-香槟分校; 亚马逊公司; 史蒂文斯理工学院; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出解耦物理建模与执行的统一框架,通过两阶段后训练策略提升小型大语言模型物理推理性能,在多个基准上实现约3%的平均性能提升。
AI 中文摘要
物理推理需要构建底层物理系统的一致模型,而非仅依赖符号或基于公式的操作。尽管大语言模型在解决数学和编码问题上表现出较强能力,但它们仍难以处理物理问题,因为这类问题将物理建模过程与数学计算交织在一起。人类处理物理问题的方式是先构建系统的表示,再进行计算。受此启发,我们提出一个统一框架,该框架提炼出明确编码物理建模过程的中间表示,并采用两阶段后训练策略:其中监督微调建立结构化建模,基于规则反馈的强化学习提升建模过程的质量。在多个多模态物理基准上的实验表明,我们的方法在不同模型和数据集上均实现了推理性能的持续提升。在PhysReason、PhyX和SeePhys基准上,物理建模输出的性能平均提升约3%,这表明显式物理建模是提升小型大语言模型物理推理能力的有效策略。
英文摘要
Physics reasoning requires constructing a consistent model of the underlying physical system rather than relying solely on symbolic or formula-based manipulation. Although large language models have shown strong ability in solving math and coding problems, they still struggle with physics problems, as these problems entangle the physical modeling process with mathematical calculations. Humans approach physics by first building a representation of the system before performing calculations. Inspired by this, we introduce a unified framework that distills intermediate representations that explicitly encode the physical modeling process and adopt a two-stage post-training strategy, where supervised fine-tuning establishes structured modeling, and reinforcement learning with rubric-based feedback improves the quality of the modeling process. Experiments on multiple multimodal physics benchmarks show that our approach generally improves physical reasoning performance across different models and datasets. Across PhysReason, PhyX, and SeePhys, physical modeling outperforms GRPO by ~3% on average. showing that explicit physical modeling is an effective strategy for improving physics reasoning in small VLMs.