追溯翻译过程:癌症细胞系中蛋白质丰度的稳定转录组预测因子的蛋白质组-wide映射
Retracing the Process of Translation: Proteome-wide mapping of stable transcriptomic predictors of protein abundance in cancer cell lines
浏览论文内容
中文总结 AI 辅助
本研究利用岭回归特征选择,在940个癌细胞系中为8,423种蛋白质识别稳定的转录组预测因子,实现可解释的蛋白质丰度建模,揭示免疫相关等生物学模块,助力假设生成与蛋白质插补。
中文摘要 AI 辅助
理解基因表达与蛋白质丰度之间的关系是分子生物学和系统生物学的核心问题。虽然基因表达反映了转录活性,但蛋白质是决定细胞表型的功能分子。然而,众多的转录后和翻译调控层使这种关系复杂化,先前的研究仅报道了RNA和蛋白质水平之间弱到中等的相关性。从转录组数据预测蛋白质丰度仍然具有挑战性,但对于生物学洞察而言,这是一个有价值的目标,尤其是在蛋白质组数据有限或不可用时。在本研究中,我们应用了一种基于岭回归的大规模特征选择策略,为940个癌症细胞系中的8,423种蛋白质识别预测性基因表达特征。据我们所知,这是首次在此规模上进行如此全面的蛋白质特异性特征选择的工作。我们的分析揭示了全局预测性和上下文特异性基因特征。这些包括具有生物学意义的模块,如免疫相关基因、HOX转录因子靶标和细胞骨架成分。模型识别出稳定的候选基因-蛋白质关联,这些关联在单个蛋白质水平和重复的转录组预测因子模式上保持可解释性。我们的方法能够从转录组数据对蛋白质表达进行可解释建模,并提供与蛋白质丰度相关的转录组特征的洞察。该框架可能支持假设生成、不完整数据集中的蛋白质插补,以及对癌症生物学中转录后调控的更深入理解。
英文摘要
Understanding the relationship between gene expression and protein abundance is central to molecular and systems biology. While gene expression reflects transcriptional activity, proteins are the functional molecules that determine cellular phenotypes. However, numerous post-transcriptional and translational regulatory layers complicate this relationship, and prior studies have reported only weak to moderate correlations between RNA and protein levels. Predicting protein abundance from transcriptomic data remains challenging, but it is a valuable goal for biological insight, especially when proteomic data is limited or unavailable. In this study, we applied a large-scale, Ridge regression-based feature selection strategy to identify predictive gene expression features for each of 8,423 proteins across 940 cancer cell lines. To our knowledge, this is the first work to perform such comprehensive protein-wise feature selection at this scale. Our analysis revealed both globally predictive and context-specific gene features. These included biologically meaningful modules such as immune-related genes, HOX transcription factor targets, and cytoskeletal components. The models identified stable candidate gene-protein associations that remained interpretable at the level of individual proteins and recurrent transcriptomic predictor patterns. Our approach enables interpretable modeling of protein expression from transcriptomic data and provides insight into transcriptomic features associated with protein abundance. This framework may support hypothesis generation, protein imputation in incomplete datasets, and deeper understanding of post-transcriptional regulation in cancer biology.
发表机构
- Universität Bielefeld(比勒费尔德大学)
机构由 AI 辅助整理,请以论文原文为准。