arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

天文基础模型中,巡天检测通道会覆盖像素并偏差层析平均红移

A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts

Ihor Kendiukhov

arXiv 2608.23626首次发表:更新:

AI 中文总结

本研究审计天文基础模型AION-1,发现巡天检测通道会系统性偏差其输出,移除该通道可消除偏差且无明显成本,偏差随模型规模增大而增强。

AI 中文摘要

天文领域的基础模型通常基于巡天像素及从这些像素衍生的星表产品进行训练,而这些星表存在可测量的不完整率,同时训练于两者的模型会将这种不完整性继承为系统性偏差。本文通过对AION-1(一个基于超过2亿个目标训练的39模态Transformer模型)的输入进行因果干预来开展审计:在保持图像令牌完全相同的情况下,仅编辑巡天分割图,就会使模型输出的所有量(流量、大小、椭率、红移)的变化幅度达到匹配安慰剂的110至4400倍。该机制属于检测门控,与场中心的存在性(r=0.47)相关,而非掩码所包含的光(r=0.30);在322个真实混合天体上,模型忽略了管道对光的划分方式(R=-0.006)。此外,这种偏好并非该通道特有:与提供所有元数据相比,矛盾的星表测光会使模型性能下降9倍。Legacy Survey管道有3.68%的目标没有覆盖其位置的片段,将该比率与管道实际返回的场所代表的缺失情况结合后,在40次分配中,层析平均红移的中位数偏移量达到LSST DESC要求的0.71倍,且在12次分配中超过该要求;观测位置误差使最差区间的偏移量达到8.3倍,按测量星等依赖性而非均匀性抽取缺失样本也不会改变这一结果。光谱学可消除该效应,而移除检测通道可在无明显成本的情况下消除该效应,且该效应随模型规模增大而增强。令牌化器存在两个进一步限制:其图像编解码器在源斑块上解析出28个有效状态,而光谱编解码器为934个;红移读出受限于量化。稀疏字典是不可靠的因果处理方式:在15次实验中,恢复率跨度为26%至75%,仅因种子差异就可产生高达18个点的变化。

英文摘要

Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic. We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports -- flux, size, ellipticity, redshift -- by 110-4400 times a matched placebo. The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all. The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times. Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale. Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑