发表机构
Universidad de Sevilla; Flatiron Institute; Institute for Advanced Study; University of Wisconsin–Madison; SLAC National Accelerator Laboratory(塞维利亚大学; 弗拉蒂伦研究所; 高等研究院; 威斯康星大学麦迪逊分校; SLAC国家加速器实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对面向物理学的机器学习中的模型误设问题,探讨其挑战、检测与缓解策略,指出需通过互补诊断迭代优化模型,强调对未知未知的稳健性源于怀疑自身模型的态度。
AI 中文摘要
机器学习目前是解决粒子物理和天文学中逆问题的核心工具,模型基于模拟数据训练并部署在真实数据上,这不仅引发了模型是否拟合的问题,还带来了模型是否以未被预期的方式出错的疑问——即“未知未知”问题。模型误设的挑战并非机器学习独有,在物理学中,误设有时正是我们想要发现的:新发现会表现为现有模型的失效;而在其他时候,我们希望这类效应被纳入分析而不扭曲测量结果。稳健的分析应能吸收我们不关注的误设,同时保留对目标误设的敏感性。机器学习既可能放大模型误设,也能提供应对它的新工具。本文讨论模型误设的挑战、检测其的诊断方法及缓解策略:没有任何单一诊断能确认模型被正确设定,检测与缓解是迭代循环的两个部分,需应用一组互补诊断、更新模型并重复该过程;对未知未知的稳健性最终不在于任何单一技术,而在于一种态度——愿意怀疑自身模型,并设计能在未被预期的错误中存活的分析方案。
英文摘要
Machine learning is now a central tool for solving inverse problems in particle physics and astronomy. Models are trained on simulation and deployed on real data, raising the question not just of whether they fit, but of whether they are wrong in ways we did not anticipate: the unknown unknowns. This challenge of model misspecification is not unique to machine learning. In physics, misspecification is sometimes exactly what we want to find: new discoveries appear as failures of existing models. At other times, we want such effects absorbed into the analysis without biasing the measurement. A robust analysis is one that absorbs the misspecifications we are not interested in, while preserving sensitivity to the ones we are. Machine learning can both amplify misspecification and provide new tools to address it. We discuss the challenges of model misspecification, diagnostics for detecting it, and strategies for mitigation. No single diagnostic can confirm that a model is correctly specified: detection and mitigation are two halves of an iterative loop, in which a battery of complementary diagnostics is applied, the model is updated, and the process repeated. Robustness against unknown unknowns is ultimately less about any single technique than about a disposition: a willingness to suspect one's own model, and to design analyses that can survive being wrong in ways one did not anticipate.
Comments25 pages, 2 figures, part of the VERaiPHY initiative