空间转录组学成本高效主动点位选择的基准测试
Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics
浏览论文内容
中文总结 AI 辅助
该研究通过回顾性基准测试,比较主动学习与随机采样在空间转录组学中的点位选择,发现主动策略在小预算下未持续优于随机采样,且性能排名因评估指标而异。
中文摘要 AI 辅助
空间转录组学(ST)在组织背景下测量基因表达,但密集的捕获网格可能成本高昂,并可能重复采样形态相似的区域。大多数主动学习策略是为分类标签和独立样本开发的。我们针对ST进行了一项基于回顾性池的主动学习与均匀随机采样的基准比较,其中表达向量是高维且连续的,候选点在空间上相关。利用两个完全剖析的公共ST队列,我们掩蔽候选表达向量,并模拟多轮选择,采用基于不确定性的蒙特卡洛dropout(MC-dropout)和时间输出差异(TOD),以及基于多样性的CoreSet和TypiClust启发的选择。我们在患者水平交叉验证下,比较了折叠范围内训练点位池的5%、10%、30%和50%处的160个完整配置,并设有独立的全标签参考。在每个预算内,策略共享选择计划、形态到表达预测器和优化协议。我们评估了每个基因的片内平均Pearson相关系数(PCC)、表达簇一致性和Moran's I保真度。在HER2阳性乳腺癌中,四种主动策略在5%、10%、30%和50%处与随机采样的汇总平均PCC差异分别为-0.0176、-0.0117、+0.0056和+0.0057。在皮肤鳞状细胞癌(cSCC)中,三种策略在5%处低于随机采样,四种策略在10%处均低于随机采样。在HER2阳性乳腺癌中,CoreSet和MC-dropout在两个最小预算下具有较低的PCC但较高的表达簇一致性;这一模式在cSCC上未重现。在报告的固定训练范围内,评估的主动策略在小预算下并未持续优于随机采样,且排名取决于评估指标。
英文摘要
Spatial transcriptomics (ST) measures gene expression in tissue context, but dense capture grids can be costly and may repeatedly sample morphologically similar regions. Most active learning strategies were developed for categorical labels and independent samples. We conduct a retrospective pool-based benchmark of active learning versus uniform Random sampling for ST, where expression vectors are high-dimensional and continuous and candidates are spatially correlated. Using two fully profiled public ST cohorts, we mask candidate expression vectors and simulate multi-round selection with uncertainty-based Monte Carlo dropout (MC-dropout) and temporal output discrepancy (TOD), and diversity-based CoreSet and TypiClust-inspired selection. We compare 160 completed configurations at 5%, 10%, 30%, and 50% of the fold-wide training spot pool under patient-level cross-validation, with a separate full-label reference. Within each budget, strategies share the selection schedule, morphology-to-expression predictor, and optimization protocol. We assess mean per-gene within-slide Pearson correlation coefficient (PCC), expression-cluster agreement, and Moran's I fidelity. On HER2-positive breast cancer, pooled mean PCC differences from Random across the four active strategies were -0.0176, -0.0117, +0.0056, and +0.0057 at 5%, 10%, 30%, and 50%, respectively. On cutaneous squamous cell carcinoma (cSCC), three strategies were below Random at 5%, and all four were below Random at 10%. On HER2-positive breast cancer, CoreSet and MC-dropout had lower PCC but higher expression-cluster agreement than Random at the two smallest budgets; this pattern did not reproduce on cSCC. Under the reported fixed training horizons, the evaluated active strategies do not consistently improve on Random at small budgets, and rankings depend on the evaluation measure.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
- Vanderbilt University(范德堡大学)
- Cornell Tech(康奈尔科技校区)
- University of Notre Dame(圣母大学)
- New York Medical College(纽约医学院)
- Vanderbilt University Medical Center(范德堡大学医学中心)
- Weill Medical College of Cornell University(康奈尔大学威尔医学院)
机构由 AI 辅助整理,请以论文原文为准。