arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

粒子物理中机器学习模式发现的分层标准:为AlphaFold时刻做准备

Stratification Criteria for Machine Learning Pattern Discovery in Particle Physics: Preparing for the AlphaFold Moment

Andrew Michael Brilliant

arXiv 2609.06015首次发表:更新:

发表机构

Applied Dynamics Research(应用动力学研究)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器学习在粒子物理中生成候选模式过快的问题,提出七项可计算的分层标准以在有限评估预算下优先排序,并建立协作式验证框架。

AI 中文摘要

机器学习能力正以加速的步伐扩展到科学领域。当应用于高能物理模式发现时,其生成的候选结果将快于传统评估所能吸收的速度。机器学习在过往数据中寻找模式,本质上具有事后性。这些模式反映的是结构还是巧合,在发现之时无从知晓;这一局限性同样适用于人类和计算模式发现。不同之处在于规模:机器学习候选生成的效率实际上是无界的,而人类评估能力保持不变。当生成速率超过评估带宽时,二元接受/拒绝便退化为随机抽样。从信息论角度看,在有限评估预算下唯一能保持排序的应对措施是分层。通过聚焦于分层而非二元过滤,规则调整可以追溯性地进行,阈值可随结果累积而调整,评估带宽可集中于排名靠前的候选。本文试图编纂这些标准,提出七项可计算评估的机器学习生成模式分层标准。目标并非给出裁决,而是优先确定哪些候选值得预先注册和纵向追踪。该框架保留了基本范式:模式加理论等于潜在的真实物理。仅凭模式本身,无论多么引人注目,在理论理解到来之前仍只是候选。明确这些标准能够在规模化预过滤的同时,创造协作资源而非竞争资源。机器学习能力扩展了物理学家可搜索的范围,同时保留了物理学家的评估方式。我们提供这一临时框架供社区校准,目标是在能力完全到来之前开发验证基础设施。

英文摘要

Machine learning capabilities are expanding into scientific domains at an accelerating pace. When applied to high-energy physics pattern discovery, they will generate candidates faster than traditional evaluation can absorb. ML finds patterns in past data it is inherently post hoc. Whether those patterns reflect structure or coincidence is unknowable at discovery time; this limitation applies equally to human and computational pattern-finding. What differs is scale: ML candidate generation is effectively unbounded, while human evaluation capacity remains fixed. When generation rate exceeds evaluation bandwidth, binary accept/reject degenerates to random sampling. Information-theoretically, the only response that preserves ranking under finite evaluation budget is stratification. By focusing on stratification rather than binary filtering, rule adjustments can be made retroactively, thresholds tuned as results accumulate, and evaluation bandwidth focused on top-ranked candidates. This paper attempts to codify those criteria, proposing seven computationally evaluable standards for stratifying ML-generated patterns. The goal is not to deliver verdicts but to prioritize which candidates merit pre-registration and longitudinal tracking. The framework preserves the essential paradigm: pattern plus theory equals potentially real physics. Patterns alone, however striking, remain candidates until theoretical understanding arrives. Making these criteria explicit enables prefiltering at scale while creating a collaborative resource rather than a competitive one. ML capabilities extend what physicists can search while preserving how physicists evaluate. We offer this provisional framework for community calibration, with the goal of developing validation infrastructure before the capability fully arrives.

Comments10 pages, 1 figure, 1 table

Journal refLondon Journal of Physics 3(1) (2026)

DOI:10.69710/ljp.v3i1.17825

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑