PhysicsBench:面向工程设计与仿真的生成式及预测式模型统一排行榜
PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation
AI总结:
PhysicsBench是面向工程设计与仿真的统一基准排行榜,涵盖多维度任务与数据集,采用标准化流程评估66个模型,可实现模型去偏排名,助力模型选择。
AI中文摘要:
生成式与预测式人工智能模型正越来越多地被用于生成几何结构,以及预测工程设计与仿真中的物理场和标量量。然而,这些模型通常是在不受约束规模的学术数据集上孤立评估的,且采用不一致的指标与流程。本文提出PhysicsBench,这是一个统一的基准与排行榜,可在单一标准化流程下评估生成式与预测式模型。PhysicsBench涵盖1D、2D和3D领域的7项生成与预测任务,对9个数据集上的66个模型进行排名,这些数据集包含工业级CAD/CFD/FEA仿真数据及公开参考数据,扩展为28种配置。生成式与预测式模型均采用同一流程与排名方式,各自在所属任务中排名;评估涵盖从S到XL的实际有限数据规模,而非学术基准中常见的无限训练集。一套通用指标套件通过分布距离衡量几何保真度,评估物理场与标量的准确性,以及工程特有的场与形状有效性。BenchRank通过两两比较优势图上的PageRank对相关指标进行去偏并排名,因此每个报告的质量指标均有对应排名,计算成本则在单独的效率视图中呈现。在所有任务中,模型的大规模学术排名与其小数据排名的相关性较弱;7项任务中有6项的顶尖模型随数据规模变化,且无模型主导超过一项任务。PhysicsBench将“最先进”从自我宣称的主张转变为公开发布的模型选择基础。
英文摘要:
Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. Yet these models are typically evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. We present PhysicsBench, a unified benchmark and leaderboard that evaluates generative and predictive models under one standardized procedure. PhysicsBench spans seven generation and prediction tasks across 1D, 2D, and 3D domains and ranks 66 models on nine datasets, comprising industrial-scale CAD/CFD/FEA simulations and public references, expanded into 28 configurations. One procedure and ranking apply to both families, each ranked within its own tasks. Evaluation spans realistic, limited data scales from S to XL rather than the unlimited training sets common in academic benchmarks. A common metric suite captures geometric fidelity with distributional distances, physical-field and scalar accuracy, and engineering-specific field- and shape-validity. BenchRank debiases correlated metrics and ranks by PageRank over a head-to-head dominance graph, so every reported quality metric is also ranked, with computational cost in a separate efficiency view. Across tasks, an architecture's large-scale academic standing weakly predicts its small-data ranking. The top model changes with data scale in six of the seven tasks, and no model leads more than one task. PhysicsBench turns "state-of-the-art" from a self-reported claim into an openly published foundation for model selection.