arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23064cs.AI

FireWorldBench:通过耦合场火灾动力学对复杂物理世界智能进行基准测试

FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics

  • Hong Kong University of Science and Technology(香港科技大学)
  • Central South University(中南大学)

机构由 AI 辅助整理,请以论文原文为准。

Qiang Chen, Hao Guo, Huatai Zhu, Tairan Huang, Yichao Cao, Hongyan Xu, Keke Huang, Haifeng Li, Yi Chen, Xiu Su

AI总结:

FireWorldBench通过耦合场火灾动力学基准,评估多模态模型和智能体在物理状态理解、因果推理及干预预测等复杂物理世界智能方面的能力,包含520个场景和9,074个问答对。

AI中文摘要:

理解物理世界需要的不仅仅是物体识别、场景描述和短期视觉预测,因为真实世界的物理系统涉及多个连续场、潜在因果机制、部分观测以及对干预敏感的动态。我们提出了FireWorldBench,一个通过耦合场火灾动力学来评估多模态大语言模型和智能体在复杂物理世界智能方面的基准。火灾提供了一个典型的压力测试环境,其中多个相互作用的物理场共同塑造可观测状态和时间动态。FireWorldBench沿两个互补的轴组织:物理能力轴和火灾场景任务轴,共同涵盖物理状态理解、时间动态、因果机制和干预推理。该基准包含520个火灾世界条目,包括494个受控模拟世界和26个真实世界对齐事件组,跨越7个环境族中的47个场景原型。这些条目结合了结构化文本观测、多个二维物理场可视化和三维事件级场景建模,产生了9,074个基于选择和开放式报告生成格式的图文交错问答对。FireWorldBench评估模型是否能够从多模态部分观测中推断潜在物理状态、解释底层机制、预测耦合场演化并评估干预后果,为复杂物理世界智能提供了一个具有挑战性的测试平台。

英文摘要:

Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields, latent causal mechanisms, partial observations, and intervention-sensitive dynamics. We propose FireWorldBench, a benchmark for evaluating complex physical world intelligence in multimodal large language models and agents through coupled-field fire dynamics. Fire provides a canonical stress-test environment, where multiple interacting physical fields jointly shape observable states and temporal dynamics. FireWorldBench is organized along two complementary axes, a physical capability axis and a fire scenario task axis, jointly covering physical-state understanding, temporal dynamics, causal mechanisms, and intervention reasoning. The benchmark comprises 520 fire-world entries, including 494 controlled simulation worlds and 26 real-world-aligned event groups, spanning 47 scene archetypes across 7 environment families. These entries combine structured textual observations, multiple 2D physical-field visualizations, and 3D event-level scene modeling, yielding 9,074 text-image interleaved question-answer pairs across choice-based and open-ended report-generation formats. FireWorldBench evaluates whether models can infer latent physical states, explain underlying mechanisms, forecast coupled-field evolution, and assess intervention consequences from multimodal partial observations, providing a challenging testbed for complex physical world intelligence.

↑