arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32857cs.CVcs.LG

H&E图像到空间转录组学是否比看起来更简单?

Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?

Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao, Hieu Pham

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过分解MSE误差结构,发现切片内变异是预测空间基因表达的关键难点,据此提出组件引导损失(CGL),其线性变体在HEST-1k上达到最优性能,证明调整训练目标比增加模型复杂度更有效。

中文摘要 AI 辅助

从常规H&E组织学预测空间基因表达,为空间分子谱分析提供了一条可扩展的途径。近期研究不断追求日益复杂的架构,以捕获空间上下文和更丰富的表达结构。与此同时,简单估计器在多项研究中展现出强劲性能,但它们已解决了哪些问题,以及何处需要额外复杂性,仍不清楚。我们通过均方误差(MSE)目标下预测误差的结构来研究这一行为。基因间平均表达的差异可解释聚合预测性能的相当大部分,而一个关键未解决的误差在于恢复每张切片内的变异。将MSE分解为切片级和切片内分量,我们发现在受控神经实验中,切片内分量具有较低的残差归一化参数敏感性。这促使我们提出组件引导损失(CGL),它增强了对切片内分量的监督。CGL-Linear是一种闭式仿射实例,在HEST-1k队列和基因面板大小上实现了整体最先进的性能。相同的切片内监督也改善了现有神经模型。这些结果表明,实质性收益可能来自将训练目标与预测误差结构对齐,而非增加模型复杂度。

英文摘要

Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in several studies, but what they already solve and where additional complexity is needed remain unclear. We study this behavior through the structure of prediction error under the mean-squared error (MSE) objective. Differences in average expression across genes can account for a substantial part of aggregate prediction performance, while a key unresolved error lies in recovering variation within each slide. Decomposing MSE into slide-level and within-slide components, we find that the within-slide component has lower residual-normalized parameter sensitivity in controlled neural experiments. This motivates Component-Guided Loss (CGL), which increases supervision of the within-slide component. CGL-Linear is a closed-form affine instantiation that achieves overall state-of-the-art performance across HEST-1k cohorts and gene-panel sizes. The same within-slide supervision improves existing neural models. These results suggest that substantial gains can come from aligning the training objective with prediction-error structure rather than increasing model complexity.

↑