arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33319cs.AI

PhysAlign:多模态物理推理中基于证据的角色对齐基准

PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning

Kecheng Liang, Haoyang Liu, Zexin Chen, Zirong Liu, Weixing Chen, Qiufeng Wang, Yang Liu, Liang Lin

首次发表
浏览论文内容

中文总结 AI 辅助

PhysAlign是一个用于评估多模态模型在物理推理中将视觉信息正确关联到物理角色的基准,通过3,341个探针和五个指标揭示了识别与角色对齐之间的显著差距。

中文摘要 AI 辅助

物理图表理解中的一个关键挑战是正确地将视觉信息与它所描述的物理实体、关系和条件关联起来。即使一个数值、符号或其他局部元素被准确识别,将其分配给错误的实体或范围也可能扭曲潜在的物理前提并导致错误的推理。为了系统地研究这一挑战,我们引入了PhysAlign,一个旨在评估多模态模型是否正确地将从物理图表中识别的信息与其预期的物理角色相关联的基准。通过局部探针和受控变体将视觉识别与物理角色分配分离,PhysAlign将对应错误与识别失败区分开来。它包含3,341个人工验证的探针,涵盖986个物理问题,使得能够在规模上系统评估视觉识别和物理角色对应。我们进一步引入了五个互补的评估指标,包括CAcc、GAcc和JAcc,它们提供了对模型识别图表内容、建立正确的物理对应关系以及解决潜在物理问题能力的全面评估。在我们评估的多模态模型中,PhysAlign揭示了局部视觉识别与物理角色基础之间的持续差距。即使查询内容被正确识别,GPT-6-Astra的条件对应错误率仍为13.8%,而InternVL3.5-8B则上升至约50.6%。这些发现表明,强大的感知能力本身并不能确保可靠的物理解释,暴露了一个主要被答案级准确性所掩盖的独特的基础瓶颈,并强调了未来模型需要更好地将识别的视觉证据与其物理意义对齐。

英文摘要

A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the underlying physical premise and lead to incorrect reasoning. To systematically study this challenge, we introduce \textbf{PhysAlign}, a benchmark designed to assess whether multimodal models correctly associate information recognized from physics diagrams with its intended physical role. By disentangling visual recognition from physical-role assignment through localized probes and controlled variants, PhysAlign isolates correspondence errors from recognition failures. It contains 3,341 human-validated probes spanning 986 physics problems, enabling systematic evaluation of visual recognition and physical-role correspondence at scale. We further introduce five complementary evaluation metrics, including CAcc, GAcc, and JAcc, which provide a comprehensive assessment of models' ability to recognize diagram content, establish correct physical correspondences, and solve the underlying physics problem. Across our evaluated multimodal models, PhysAlign reveals a consistent gap between local visual recognition and physical-role grounding. Even when the queried content is correctly recognized, the conditional correspondence error rate remains 13.8\% for GPT-6-Astra and rises to about 50.6\% for InternVL3.5-8B. These findings indicate that strong perception alone does not ensure reliable physical interpretation, exposing a distinct grounding bottleneck that is largely hidden by answer-level accuracy and highlighting the need for future models to better align recognized visual evidence with its physical meaning.

↑