arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11185cs.LGq-bio.GN

细胞命运分配中的预测多重性:无标签Rashomon集与单细胞认证的局限性

Predictive Multiplicity in Cell-Fate Assignment: Label-Free Rashomon Sets and the Limits of Per-Cell Certification

Arjun Bhupatiraju, Abhiram Bhupatiraju

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对单细胞轨迹推断中同等拟合数据的模型配置会分配冲突细胞命运的问题,提出无标签框架FateMultiplicity构建Rashomon集,发现单细胞认证的命运边际FM无法提升分配可靠性,仅边际侵蚀比等两种构建方法有效。

中文摘要 AI 辅助

单细胞轨迹推断将转录组学测量结果映射到发育连续体上,但能同等拟合数据的配置可能会分配相互冲突的细胞命运。FateMultiplicity是一种无标签框架,它在无需谱系标签的情况下构建统计上可接受的模型集(即Rashomon集),方法是在针对随机种子变异校准的非劣效性检验下,评估交叉拟合的保留基因上的模型差异。多重性较大,且更多取决于模型空间的多样性而非其规模:第二种算法的12种配置揭示了20.0%的细胞,而第一种算法的24种配置仅揭示了3.8%的细胞。随后测试了单细胞认证的命运边际FM是否能比拟合模型本身提供更可靠的分配,结果显示其并不能。在模拟真实值下,针对相同细胞,FM区分错误分配的AUC为0.682,而基线配置自身的决策边际AUC为0.965(p=0.003),种子分散基线的AUC为0.854。信息性由可接受集的广度而非基数决定:在基数为4时,种子重拟合得到的AUC为0.933,超参数扰动集的AUC为0.701。将下确界放宽至q分位数可恢复区分能力,但会收敛至单个模型自身的置信度;上确界达到0.973,因为θ*的成员资格从下方对其进行约束,而下确界无锚定。轨迹推断中的多重性值得测量和报告,但在无标签Rashomon集上进行单细胞认证并非获得更可靠命运调用的途径。两种构建方法有效:边际侵蚀比在模拟中区分真实与虚假分支点(AUC=0.890,未在真实数据上测试);针对克隆观察到的命运,未认证细胞与其克隆结果的不一致率比认证细胞高16.4个百分点(p<0.001)。

英文摘要

Single-cell trajectory inference maps transcriptomic measurements onto developmental continua, yet configurations that fit the data equally well can assign conflicting cell fates. FateMultiplicity is a label-free framework that constructs a statistically admissible model set, or Rashomon set, without lineage labels, by evaluating model discrepancy on cross-fitted held-out genes under non-inferiority testing calibrated against random-seed variation. Multiplicity is large and depends more on the diversity of the model space than its size: twelve configurations of a second algorithm expose 20.0% of cells where twenty-four of the first expose 3.8%. Whether the per-cell certified fate margin FM yields more reliable assignments than the fitted model already provides is then tested, and it does not. On simulation ground truth, on the same cells, FM discriminates misassignment at AUC 0.682, against 0.965 for the baseline configuration's own decision margin (p = 0.003) and 0.854 for a seed-dispersion baseline. Informativeness is governed by the breadth of the admitted set, not its cardinality: at cardinality four, seed refits give 0.933 and hyperparameter-perturbed sets 0.701. Relaxing the infimum to a q-quantile recovers discrimination but converges toward the single model's own confidence; the supremum reaches 0.973 because theta*'s membership bounds it from below, while the infimum is unanchored. Multiplicity in trajectory inference is worth measuring and reporting, but per-cell certification over a label-free Rashomon set is not a route to more reliable fate calls. Two constructions survive: a margin-erosion ratio separates real from spurious branch points in simulation (AUC 0.890, untested on real data), and against clonally observed fate, uncertified cells disagree with their clone's outcome 16.4 percentage points more often than certified cells (p < 0.001).

发表机构

  • McNeil High School(麦克尼尔高中)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑