发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对钙群体动力学模型跨数据集和物种推广难的问题,提出CAPT这一连续自回归群体Transformer,通过连续补丁令牌化策略建模,经多数据集实验验证,其在预测任务中性能优且能形成共享功能空间,为通用神经基础模型开辟新途径。
AI 中文摘要
大规模钙成像为构建神经群体动力学的基础模型创造了机会,但一个核心问题仍未解决:在一组记录上预训练的模型能否推广到新数据集、实验范式甚至物种。现有方法通常针对特定任务设计并在单个数据集上评估,其学习表示能否用于新钙痕数据集尚不清楚。为解决这一差距,我们提出了CAPT,一种用于钙群体动力学的连续自回归群体Transformer。CAPT通过连续补丁令牌化策略直接对连续钙痕建模并自回归训练,实现端到端预训练和适应各种下游任务。我们首先在大规模小鼠钙成像数据集上预训练CAPT,并评估其在不同实验室收集的独立小鼠、斑马鱼幼虫和秀丽隐杆线虫数据集上的可转移性。在这些转移设置中,预训练的主干被冻结,只更新适应模块。在神经群体预测和行为解码任务中,CAPT始终优于专门和通用的基线。除了预测性能,使用秀丽隐杆线虫数据集的NeuroPAL注释进行的多模态分析表明,CAPT嵌入在数据集之间形成了一个共享功能空间,并捕获了解剖细胞身份相关结构。这些结果表明,连续自回归建模为钙成像通用神经基础模型开辟了一条简单途径,该模型可跨数据集、实验范式和物种推广。
英文摘要
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species.} Existing approaches are often designed for specific tasks and evaluated on a single dataset, making it unclear whether their learned representations are reusable for new calcium trace datasets. To tackle this gap, we present \textbf{CAPT}, a \textbf{C}ontinuous \textbf{A}utoregressive \textbf{P}opulation \textbf{T}ransformer for calcium population dynamics. CAPT models continuous calcium traces directly through a continuous patch tokenization strategy and is trained autoregressively, enabling end-to-end pretraining and adaptation to diverse downstream tasks. We first pretrain CAPT on a large-scale mouse calcium imaging dataset and evaluate its transferability across independent mouse, larval zebrafish, and \textit{C. elegans} datasets collected by different laboratories. In these transfer settings, the pretrained backbone is frozen and only adaptation modules are updated. Across neural population forecasting and behavior decoding tasks, CAPT consistently outperforms specialized and general-purpose baselines. Alongside predictive performance, multimodal analyses using NeuroPAL annotations in \textit{C. elegans} datasets show that CAPT embeddings form a shared functional space across datasets and capture anatomical cell-identity-related structure. These results suggest that the continuous autoregressive modeling opens up possibilities for a simple route towards general-purpose neural foundation models for calcium imaging, which can generalize across datasets, experimental paradigms, and species. Code is available at https://github.com/TSuXinH/CAPT.