arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

发现物理表示语言

Discovering Physical Representation Languages

Linzhe Zhang, Changming Xu

arXiv 2609.23381首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出物理表示语言发现,通过匿名受控实验恢复隐藏本体论,给出可辨识性理论、多项式时间算法及极小极大界,将科学ML从学习定律转向发现表达定律的语言。

AI 中文摘要

在机器能够发现物理定律之前,它必须发现其测量值是什么:哪些观测位于胞腔上,哪些是强度量或广延量,哪些扇区是对偶的,哪些区分仅仅是规范。我们引入了物理表示语言发现这一课题,即直接从匿名受控实验中恢复这一隐藏本体论的问题。我们给出一个可辨识性理论和构造性多项式时间过程,该过程能恢复载体和微分序列、测量类型和方向扭转、不可逆细化语义、原始-对偶麦克斯韦图,以及任何允许的实验都无法打破的残余等价关系。该理论将物质干扰转化为交换子,利用细化将量与坐标分离,并且仅在表示被恢复之后才选择物理。对于经过认证的有限实验族,我们证明了端到端的两阶段测量界以及在维度、精度和置信度上的匹配极小极大速率。盲麦克斯韦实验在联合损坏的观测下,在规则和不规则载体上恢复了完整的原始/相对对偶本体论;一个独立的不规则RLC系统表明该结果并非麦克斯韦所特有。该框架可扩展到每个载体数万个胞腔,而压力审计展示了在严酷物理状态下的鲁棒性——包括非马尔可夫记忆、非线性、非局域性和复杂本构滞后。一项公开的FDTD审计展示了从不完整场数据中涌现的匿名旋度结构,同时刻画了完整恢复所需的信息前提。目标是推动科学机器学习从在人类提供的语言中学习定律,转变为发现定律得以表达的语言,从而确立观测可辨识性的精确理论极限。

英文摘要

Before a machine can discover a physical law, it must discover what its measurements are: which observations live on cells, which are intensive or extensive, which sectors are dual, and which distinctions are merely gauge. We introduce physical representation-language discovery, the problem of recovering this hidden ontology directly from anonymous controlled experiments. We give an identifiability theory and constructive polynomial-time procedure that recovers a carrier and differential sequence, measurement types and orientation twist, noninvertible refinement semantics, primal-dual Maxwell diagrams, and the residual equivalences that no permitted experiment can break. The theory turns material nuisance into a commutant, uses refinement to separate quantities from coordinates, and selects physics only after its representation has been recovered. For a certified finite experiment family, we prove an end-to-end two-stage measurement bound and a matching minimax rate in dimension, accuracy, and confidence. Blind Maxwell experiments recover complete primal/relative-dual ontologies on regular and unstructured carriers under jointly corrupted observations; an independent unstructured RLC system demonstrates that the result is not specific to Maxwell. The framework scales to tens of thousands of cells per carrier, while stress audits demonstrate robustness across severe physical regimes - including non-Markovian memory, nonlinearities, nonlocality, and complex constitutive hysteresis. A public FDTD audit demonstrates the emergence of anonymous curl structure from incomplete field data, while characterizing the informational prerequisites for complete recovery. The goal is to move scientific ML from learning laws in a human-supplied language to discovering the language in which laws become expressible, establishing exact theoretical limits on observational identifiability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑