arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索物理问题中的结构:AI智能体能否发现统计力学映射?

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Wanyu Zhao, Wanbing Zhao

arXiv 2607.26367首次发表:更新:

AI 中文总结

本研究以StatMechBench-v0为基准,评估基于LLM的“提出-验证-修正”智能体发现统计力学映射的能力,揭示其推理局限并提出验证栈的设计方向。

AI 中文摘要

理论物理学的一项重要技能是识别新问题何时可转化为已知模型,我们将该技能作为AI智能体任务进行研究:基于大语言模型(LLM)的智能体能否从原始配分函数中发现统计力学映射,将其转化为可处理的表示?为探究该问题,我们推出StatMechBench-v0,这是包含6个伊辛型问题的基准,涵盖转移矩阵方法、规范可移除无序、平面/Pfaffian结构。我们在多个LLM及问题表述下评估简单的“提出-验证-修正”智能体,结果显示数值反馈常帮助智能体修复代码并恢复正确配分函数,但智能体也可能通过数值检查却错误识别底层可处理类别或低估计算复杂度,这既揭示了当前LLM推理的局限,也要求构建超越数值一致性的验证栈,例如纳入符号检查和结构不变量。本研究为针对理论物理结构发现的AI智能体提供了早期评估及设计方向。

英文摘要

An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representation? To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and planar/Pfaffian structure. We evaluate a simple propose-verify-revise agent across multiple LLMs and problem phrasings. The results show that numerical feedback often helps agents repair code and recover correct partition functions. However, agents can also pass the numerical checks while misidentifying the underlying tractable class or understating computational complexity. This both reveals limitations in current LLM reasoning and calls for a verification stack that goes beyond numerical agreement, incorporating, for example, symbolic checks and structural invariants. Our study provides an early evaluation and design directions for AI agents aimed at structural discovery in theoretical physics.

CommentsAccepted to the AID-Wild workshop at CAIS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑