发表机构
Rensselaer Polytechnic Institute; University of California Santa Cruz(伦斯勒理工学院; 加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对高密度微电极阵列的稀疏放电数据,提出基于时空基元词汇表的离散生成模型,在31项人脑类器官与海马组织实验中实现5.2倍重建精度提升,并建立可复用的阵列级表示。
AI 中文摘要
神经活动的生成模型有助于表征组织动态、比较实验条件,并模拟群体活动,其应用范围涵盖疾病与药物反应研究以及闭环实验。然而,现有方法通常假设存在一组固定的已分选神经元,而高密度微电极阵列产生的则是极其稀疏、覆盖整个阵列的二元放电数据体,其中观察到的电极子集会因实验而异。我们引入了一种离散生成模型,利用共享的时空基元词汇表来表示这种活动。残差向量量化自编码器学习基元词汇表,而因子化掩码变换器则预测活动发生的位置以及每个活动位置出现的基元。我们在涵盖人脑类器官和急性离体人海马组织的31项实验中评估了该模型。学习到的基元被广泛复用:实验身份仅解释了基元使用熵的9%,且不同组织类型间的基元重叠程度与组织内部的基元重叠程度相当。当独立于生成先验评估表示质量时,我们的方法在体素级重建平均精度上达到了匹配的扁平分词器的5.2倍。对于掩码补全和自由生成,完整模型在点位级平均精度上达到了匹配的生成基线的1.4至2.6倍,并在所有四类生成指标上均优于该基线。这些结果为阵列级放电活动建立了一种紧凑且可复用的表示,无需学习特定于实验的参数,为跨多种神经制备的生成建模提供了可扩展的基础。
英文摘要
Generative models of neural activity could help characterize tissue dynamics, compare experimental conditions, and simulate population activity for applications ranging from disease and drug-response studies to closed-loop experimentation. Existing approaches, however, typically assume a fixed set of sorted neurons, whereas high-density microelectrode arrays produce extremely sparse, array-wide binary spike volumes in which the observed subset of electrodes varies across assays. We introduce a discrete generative model that represents this activity using a shared vocabulary of spatiotemporal motifs. A residual vector-quantized autoencoder learns the motif vocabulary, while a factorized masked transformer predicts where activity occurs and which motif appears at each active location. We evaluate the model on 31 assays spanning human brain organoids and acute \emph{ex vivo} human hippocampal tissue. The learned motifs are broadly reused: assay identity explains only $9%$ of the entropy in motif use, and motif overlap across tissue types is comparable to overlap within them. When representation quality is evaluated independently of the generative prior, our approach achieves $5.2\times$ the voxel-level reconstruction average precision of a matched flat tokenizer. For masked completion and free generation, the full model achieves $1.4$--$2.6\times$ the site-level average precision of the matched generative baseline and outperforms it across all four families of generation metrics. These results establish a compact, reusable representation for array-wide spiking activity without learned assay-specific parameters, providing a scalable foundation for generative modeling across diverse neural preparations.