arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31201cs.AI

样本、来源、空间:在人类大脑微结构空间结构化表征学习中分解数据规模

Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将数据规模分解为样本数量、来源多样性和空间覆盖度,在人类大脑微结构表征学习中验证了这些轴对性能与泛化的不同影响。

中文摘要 AI 辅助

规模研究通常将训练数据表示为单一数量的样本。然而,对于层次化和空间结构化的数据,相同数量的样本可以从少量或大量来源中抽取,并分布在不同底层域中。因此,我们将数据规模研究视为一个分配问题,将唯一样本数量、来源多样性和空间覆盖度分开考虑。我们在微观全脑组织学中研究这种分解,其中来源是单个大脑,样本是特定空间位置的图像块。在93次受控预训练运行中,使用空间邻近性作为监督的对比模型,我们在来自21个人类大脑的1160万个空间锚定图像块上变化数据分配、计算量和模型容量。性能随着更多唯一样本、更广空间覆盖、额外计算量和更大模型容量而提高。在固定样本数量下,将样本分布在1到18个受试者之间未产生可检测的改善,尽管表征对预训练中遇到的受试者泛化效果显著更好。因此,受试者间差异强烈影响泛化,但当固定样本预算分布在更多来源时,额外受试者不提供益处。这些结果确立了样本数量、来源多样性和空间覆盖度作为空间结构化表征学习中数据规模的独立轴。

英文摘要

Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.

发表机构

  • Research Centre Jülich(于利希研究中心)
  • Helmholtz AI(亥姆霍兹人工智能)
  • Jülich Supercomputing Centre(于利希超算中心)
  • Frankfurt Institute for Advanced Studies(法兰克福高等研究院)
  • Goethe University Frankfurt(法兰克福大学)
  • University Hospital Düsseldorf(杜塞尔多夫大学医院)
  • University of Koblenz(科布伦茨大学)

机构由 AI 辅助整理,请以论文原文为准。

↑