AI 中文总结
本章介绍面向天文领域的基础模型,阐述其核心是可迁移表示,指出当前天文领域此类模型的可迁移性验证案例较少,未来需进一步探索提升路径。
AI 中文摘要
基础模型是在广泛数据上预训练一次,随后可复用至众多任务的高容量网络。本章面向天文研究者,从可迁移表示的概念切入展开介绍,可迁移表示是网络在训练过程中形成的内部描述,而非拟合的任务本身,才是可迁移至新问题的核心。我们从基础原理出发,先阐述表示的重要性及有用表示的特征,再梳理构建这类表示的架构、自监督目标、缩放方法、适配技术及跨模态学习。贯穿全文的核心是区分这些方法与它们服务的目标:拥有Transformer、自监督目标及大规模预训练本身并不足以让模型成为基础模型,其核心属性是学习到的表示具备可迁移性,可通过仅用少量或无任务特定训练数据就能在新任务上工作(少样本和零样本学习)的能力来验证。接着我们聚焦天文学领域,该领域数据充足但标签稀缺,且常用模拟数据替代真实值。我们对当前相关文献持审慎态度:许多模型采用了基础模型的架构,但在仪器、天体群体及任务间实现清晰可迁移的案例相对较少。这一现象符合预期,因为即使在计算机视觉及更广的物理科学领域,语言之外的稳健可迁移性仍不常见,进一步缩放模型或采用不同的表示理论能否缩小这一差距仍是待解决的问题。最后我们将该目标置于更广泛的机器智能目标框架下,并概述了标志着真正进步所需的证据。
英文摘要
Foundation models are high-capacity networks pretrained once on broad data and then reused across many tasks. This chapter introduces them through the idea of a transferable representation, the internal description a network forms during training, which, rather than the fitted task, is what carries over to new problems. We develop the idea from first principles for an astronomical reader, starting from why a representation matters and what makes one useful, and then surveying the architectures, self-supervised objectives, scaling, adaptation, and cross-modal learning that produce one. A theme throughout is the distinction between these methods and the goal they serve. The presence of a transformer, a self-supervised objective, and large-scale pretraining does not by itself make a model a foundation model, since the defining property is that the learned representation transfers, as tested by its ability to work on new tasks with little or no task-specific training data (few-shot and zero-shot learning). We then consider astronomy, where data are abundant but labels are scarce and simulations often stand in for ground truth. Here we offer a cautious reading of the current literature, in which many models adopt the architecture of foundation models while clear demonstrations of transfer across instruments, populations, and tasks remain comparatively rare. This is to be expected, since robust transfer beyond language is still uncommon even in vision and the wider physical sciences, and whether further scaling or a different account of representation will close the gap remains an open question. We close by placing the goal within the broader aim of machine intelligence and outlining the evidence that would mark real progress.
CommentsInvited chapter for the edited book "Machine Learning Techniques for Astrophysics and Cosmology" (Eds. Cosimo Bambi, Vinay Kashyap, Swarnim Shashank, Naoki Yoshida, Springer Singapore, expected in 2026). Submitted version