CRNDiff:基于化学反应网络的计数原生扩散框架
CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks
浏览论文内容
中文总结 AI 辅助
CRNDiff提出基于化学反应网络的计数原生扩散框架,利用闭式转移核和倾斜Feynman-Kac转向,在单细胞数据上实现高保真条件生成,优于现有模型。
中文摘要 AI 辅助
诸如单细胞RNA(scRNA)测序等科学测量通常以非负整数计数的形式呈现,而连续状态扩散模型则使用连续坐标来近似这种离散结构。基于随机化学反应网络(CRNs)——一类计数原生的马尔可夫跳跃过程,我们引入了CRNDiff,这是一个结构化框架,它将计数空间扩散与对稀有亚群的推理时条件化相结合。一个独立的生灭实例为前向加噪提供了闭式转移核。该核使得通过前向滤波-后向采样(FFBS)进行反向采样成为可能,并支持数据驱动的终端加噪时间选择,从而无需验证集扫描。这种可处理性还使我们能够引入倾斜的Feynman-Kac(FK)转向方法,该方法无需重新训练即可从冻结的生成器中采样目标亚群。通过在FK粒子校正之前倾斜后验边际,转向方法减轻了当目标群体稀有时重要性权重的集中问题。利用人类心脏细胞图谱的scRNA-seq数据,我们测试了CRNDiff生成细胞类型特异性分布的能力。在评估的三个目标群体中,CRNDiff在评估的生成模型中实现了最高的条件保真度,且对于更稀有的目标群体具有更大的平均纯度边际。生成的细胞保留了标记水平的差异表达结构。用生成的细胞替换目标类别的真实训练细胞,所得的下游分类性能接近真实数据参考。
英文摘要
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning on rare subpopulations. An independent birth--death instantiation yields a closed-form transition kernel for forward noising. This kernel enables reverse sampling via forward-filtering backward-sampling (FFBS) and supports data-driven selection of the terminal noising time, eliminating the need for a validation sweep. This tractability also lets us introduce tilted Feynman--Kac (FK) steering, a method for sampling target subpopulations from a frozen generator without retraining. By tilting posterior marginals before FK particle correction, steering mitigates importance-weight concentration when the target population is rare. Using scRNA-seq data from the human heart cell atlas, we test the ability of CRNDiff to generate cell-type-specific distributions. Across the three evaluated target populations, CRNDiff achieves the highest conditional fidelity among the evaluated generative models, with larger mean purity margins for rarer target populations. Generated cells preserve marker-level differential-expression structure. Replacing real training cells for the target classes with generated cells yields downstream classification performance approaching that of the real-data reference.
发表机构
- Institute of Industrial Science(工业科学研究所)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。