发表机构
Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对转录组数据的获取限制,提出并基准测试三种生成对抗网络变体,其中MK-TGAN模型结合图神经网络整合先验知识,生成的合成转录组数据真实性与实用性更优。
AI 中文摘要
随着生物医学研究越来越依赖数据密集型工具,数据集的质量和实用性至关重要。数据不平衡、偏差以及伦理或法律约束等挑战往往限制了对高质量数据的获取。合成数据生成可帮助克服这些局限。本文对转录组数据的生成模型开展对比分析,研究通过基因图整合先验生物学知识的策略,确保合成数据能捕捉真实世界的基因模式,保持其对下游任务的实用性。特别地,本文引入并基准测试了生成对抗网络(GAN)的三个变体。在这些备选模型中,MK-TGAN——一种创新性多内核、基于图神经网络(GNN)的模型——在生成数据的真实性和实用性方面的表现尤为突出。与其他方法不同,MK-TGAN通过利用图神经网络来利用先验知识图。研究结果表明,先验知识整合策略可提升性能,且MK-TGAN始终能生成具有更优真实性和生物学合理性的合成样本。
英文摘要
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.