AI 中文总结
该研究针对物理学中域适应的假设不成立问题,提出自适应域适应方法,通过重加权模拟事件聚焦物理不匹配,并给出无标签模型选择规则以选择近最优操作点。
AI 中文摘要
域适应被广泛用于使在模拟数据上训练的神经网络适用于实验数据,其前提是两个域仅在干扰因素上存在差异,且关注量在两个域中的分布完全相同。但在物理学中,这两个假设均不成立:模拟可能对物理规律的描述有误,而目标关注量(如能谱、红移分布)的分布往往就是测量本身。我们在一个空气簇射玩具基准测试中研究了这类不匹配的后果,该基准测试中探测器响应干扰、物理模拟偏移和能谱偏移可单独或共同开启。标准对抗性域适应可处理条件偏移,但一旦两个能谱不同,它会对齐两者,将不受控制的偏差替换为锚定在模拟先验上的偏差。我们提出自适应域适应,该方法对模拟事件进行重加权,以使域适应仅聚焦于真正的物理不匹配。由于预测的能谱依赖于模型训练配置,我们提供了一种无标签的模型选择规则,用于选择接近最优的操作点。
英文摘要
Domain adaptation is widely used to make neural networks trained on simulations applicable to experimental data. Its premise is that the two domains differ only in nuisances, and that the quantity of interest is distributed identically in both. In physics neither assumption holds: simulations can be wrong about the physics, and the distribution of the target quantity - an energy spectrum, a redshift distribution - is often the measurement itself. We study the consequences of such mismatches on a toy air-shower benchmark in which a detector-response nuisance, a physical simulation shift, and an energy-spectrum shift can be switched on separately or together. Standard adversarial adaptation handles the conditional shifts, but once the two spectra differ it aligns them, replacing an uncontrolled bias by one anchored on the simulation prior. We present adaptive domain adaptation, which reweights the simulated events so as to focus domain adaptation on the genuine physical mismatch alone. Since the predicted spectrum depends on model training configuration, we provide a label-free model selection rule for selecting the near-the-best operation point.