arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

星系目录场级宇宙学推断中的归纳偏置

Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs

James O. Baldwin, Shy Genel, Francisco Villaescusa-Navarro

arXiv 2609.09504首次发表:更新:

发表机构

City University of New York Graduate Center; Center for Computational Astrophysics, Flatiron Institute; Columbia University; Princeton University(纽约城市大学研究生中心; 熨斗研究所计算天体物理学中心; 哥伦比亚大学; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过比较Deep Sets与图神经网络在CAMELS模拟星系目录中的推断性能,发现本动速度是集合模型获取$\Omega_m$信息的主要来源,而空间关系需显式编码才能有效利用。

AI 中文摘要

我们使用具有不同归纳偏置的机器学习模型,从模拟星系目录中对物质密度参数$\Omega_m$进行场级似然无关推断。利用CAMELS的流体动力学模拟,我们考察了观测选择和架构如何控制宇宙学信息提取。我们分别及联合考虑星系位置和视线方向本动速度,并将置换不变的Deep Sets(分别用多层感知机(MLP)或Kolmogorov-Arnold网络(KAN)实现)与显式编码空间关系的图神经网络(GNN)进行比较。我们测试了具有不同亚网格星系形成处方模拟中的分布内和分布外(OOD)性能。Deep Sets仅从速度推断$\Omega_m$,分布内平均相对误差约为18%,OOD约为25%,KAN和MLP性能相当。相反,相同的基于集合的方法在分布内或跨套件测试中均未产生有用的$\sigma_8$预测。添加位置信息并未改善Deep Sets,而GNN推断$\Omega_m$的分布内平均相对误差约为10%,OOD为10%–17%。这些结果表明,在此设置中,本动速度为基于集合的模型提供了$\Omega_m$信息的主要来源,而空间信息最有效地被显式编码星系-星系关系的架构所利用。由于速度输入是精确的模拟本动速度,应用于巡天数据需要在真实的速度测量噪声、选择效应和巡天几何下进行验证。

英文摘要

We perform field-level likelihood-free inference of the matter density parameter $Ω_m$ from simulated galaxy catalogs using machine learning models with differing inductive biases. Using hydrodynamic simulations from CAMELS, we examine how observable choice and architecture govern cosmological information extraction. We consider galaxy positions and line-of-sight peculiar velocities, separately and jointly, and compare permutation-invariant Deep Sets, implemented with either multilayer perceptrons (MLPs) or Kolmogorov-Arnold Networks (KANs), to graph neural networks (GNNs), which explicitly encode spatial relations. We test in-distribution and out-of-distribution (OOD) performance across simulations with different subgrid galaxy-formation prescriptions. Deep Sets infer $Ω_m$ from velocities alone with mean relative errors of approximately $18\%$ in-distribution and $\sim25\%$ OOD, with KANs and MLPs achieving comparable performance. In contrast, the same set-based approach does not yield useful $σ_8$ predictions in either in-distribution or cross-suite tests. Adding positions does not improve Deep Sets, while GNNs infer $Ω_m$ with mean relative errors of about $10\%$ in-distribution and $10$--$17\%$ OOD. These results indicate that peculiar velocities provide the dominant source of $Ω_m$ information for set-based models in this setting, while spatial information is most effectively used by architectures that explicitly encode galaxy-galaxy relations. Because the velocity inputs are exact simulated peculiar velocities, applications to survey data will require validation under realistic velocity-measurement noise, selection effects, and survey geometry.

Comments23 pages, 8 figures, 5 tables. Accepted for publication in The Astrophysical Journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑