AI 中文总结
ALICE是一个基础模型,通过上下文内估计整流流速度场实现零样本互信息估计,无需逐分布训练,在多个科学领域达到与专用神经估计器相当的精度。
AI 中文摘要
从样本中估计互信息(MI)是众多科学领域的核心目标。现代神经估计器在大数据场景下表现准确,但在数据稀缺时则有所不足,且对于每个研究的分布都必须重新拟合。此外,当前的估计器还受限于特定的数据类型。这些限制阻碍了它们在许多应用中的采用,在这些应用中,针对每个分布的训练不切实际,且样本量较小。我们提出了ALICE,一个消除了逐分布训练的基础模型,同时实现了具有竞争力的估计精度。ALICE仅在广泛的合成分布族上训练,充当整流流速度场的上下文内估计器:以未见分布的样本为条件,它无需任何显式训练即可估计该分布的速度场。随后,通过一个固定恒等式获得互信息,该恒等式整合了联合场与条件场之间的平方差。我们在一个标准且具有挑战性的基准上验证了ALICE,并将其应用于三个领域:生物学、遗传学和神经科学,这些领域的数据模型从未见过。我们首次证明,单个模型缩小了与为每个分布单独训练的神经估计器之间的差距,同时原生支持不同的数据维度和样本数量,实现了跨科学领域的零样本互信息分析。
英文摘要
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are moreover tied to specific data types. These constraints limit their adoption in many applications where per-distribution training is impractical and sample sizes are small. We present ALICE, a foundation model that removes per-distribution training, while achieving competitive estimation accuracy. Trained exclusively on a broad family of synthetic distributions, ALICE acts as an in-context estimator of rectified-flow velocity fields: conditioned on samples of an unseen distribution, it estimates that distribution's velocity field without any explicit training. MI is then obtained through a fixed identity that integrates the squared difference between the joint and conditional fields. We validate ALICE on a standard, challenging benchmark and apply it in three domains, biology, genetics, and neuroscience, whose data the model has never seen. For the first time, we show that a single model closes the gap with neural estimators trained separately for each distribution, while natively supporting different data dimensionality and sample cardinality, enabling zero-shot MI analysis across scientific domains.