arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

社区中心在高维、小样本基因表达数据中识别稳健生物标志物

Community-Centers Identify Robust Biomarker in High-Dimensional, Low-Sample-Size Gene Expression Data

Ruiqi Li, Te Bai, Renaud Lambiotte, Orr Levy, Paul Expert, Daqing Li, Shlomo Havlin

arXiv 2610.00986首次发表:更新:

发表机构

Beihang Valencia Polytechnic Institute, Beihang University; Hangzhou International Innovation Institute, Beihang University; Shanghai Key Laboratory of Intelligent Information Processing, Fudan University; Mathematical Institute, University of Oxford; Faculty of Engineering, Bar-Ilan University; Global Business School for Health, Faculty of Population Health Sciences, University College London; College of Safety Science and Engineering, Civil Aviation University of China; Department of Physics, Bar-Ilan University(北京航空航天大学; 北京航空航天大学杭州创新研究院; 复旦大学; 牛津大学数学研究所; 巴伊兰大学工程学院; 伦敦大学学院人口健康科学学院全球健康商学院; 中国民航大学安全科学与工程学院; 巴伊兰大学物理系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维小样本基因表达数据,提出社区中心回归(RCC)框架,利用网络相变识别稳健生物标志物,在五个数据集上显著优于基准方法。

AI 中文摘要

高维、小样本的批量基因表达数据对转录组学构成了根本性挑战,这类数据通常包含数万个基因但样本数量相对较少,导致过拟合以及关键特征(作为生物标志物)选择的不稳定。我们提出了一种针对此类数据的社区中心回归(RCC)框架。RCC将基因表达数据转换为特征邻近网络,并利用巨连通分量的相变来识别一个临界距离阈值,该阈值能够产生一个稀疏但信息量最大的网络表示,其标志是最小归一化最短压缩长度。直观上,在临界状态下,网络在碎片化与过度连接之间取得平衡,使得有意义的基因-基因关联社区得以涌现,同时过滤噪声,从而让社区中心捕获最具信息量且非冗余的信号。这些代表性基因随后被输入一个简单的普通最小二乘(OLS)模型进行下游预测。在五个批量基因表达数据集上,RCC展现出卓越的预测性能,明显优于基准方法,对缺失数据和噪声保持稳健,且不依赖外部生物学知识。这些结果表明,基于网络的表示为转录组学之外的高维、小样本预测任务提供了一种有效且通用的框架,并展示了网络科学思想如何支持数据稀缺场景下的稳健学习。

英文摘要

High-dimensional, low-sample-size bulk gene expression data poses a fundamental challenge in transcriptomics, which typically includes tens of thousands of genes but relatively few samples, leading to overfitting and unstable selection for key features as biomarkers. We propose a regression by community centers (RCC) framework tailored for such data. RCC converts gene expression data into a feature proximity network and leverages the phase transition of the giant connected component to identify a critical distance threshold that yields a sparse yet maximally informative network representation, indicated by a minimum normalized shortest compression length. Intuitively, at the criticality, the network balances fragmentation and over-connectivity, allowing meaningful gene-gene associations communities to emerge while filtering noises, such that community-centers capture the most informative and non-redundant signals. These representative genes are then fed to a simple ordinary least squares (OLS) model for downstream predictions. Across five bulk gene expression datasets, RCC displays exceptional prediction performance, clearly outperforming benchmark methods, remains robust to missing data and noises, and does not rely on external biological knowledge. These results suggest that network-based representations provide an effective and general framework for high-dimensional, low-sample-size prediction tasks beyond transcriptomics and illustrate how network-science ideas can support robust learning in data-scarce regimes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑