发表机构
National Institute of Research and Development for Biological Sciences; University of Turku; Research Institute for Artificial Intelligence “Mihai Drăgănescu”, Romanian Academy; University of Bucharest(罗马尼亚生物科学研究与发展国家研究所; 图尔库大学; 罗马尼亚科学院“米哈伊·德拉格内斯库”人工智能研究所; 布加勒斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对脑胶质瘤拉曼光谱数据集小且异质的问题,开发条件变分自编码器生成合成光谱增强真实训练数据,在严格协议下评估,结果表明该方法能提升分类性能,支持深度生成式增强可提高机器学习在相关应用中的鲁棒性。
AI 中文摘要
获取足够大的生物医学数据集仍然是基于拉曼光谱诊断的机器学习的主要障碍。特别是对于脑胶质瘤分析,数据集通常小且异质,受采集特定变异性影响。本文研究了在小样本队列中深度生成式增强的效用。分析了从58个肿瘤样本获取的脑胶质瘤活检光谱,考虑二元IDH状态分类和6类甲基化亚型分类问题。为解决数据集规模有限和不平衡问题,开发了能生成类条件合成拉曼光谱的条件变分自编码器(β-CVAE)。在严格的患者隔离交叉验证协议下,在Train-on-Synthetic, Test-on-Real(TS/TR)和Train-on-Synthetic+Real, Test-on-Real(TSR/TR)设置中评估生成的数据。仅在合成数据上训练的模型表现不如在真实光谱上训练的模型,表明合成与真实分布之间存在显著域差距。然而,用合成光谱增强真实训练数据持续提高了多个模型的分类性能。这些发现表明,即使独立患者样本数量有限,生成模型也能捕获足够结构为下游分类器提供有用正则化。还研究了一种基于重建的推理策略,即通过重建分类(CbR),其中类预测基于不同类条件下的重建误差。总体而言,结果支持使用深度生成式增强作为一种实用策略,以提高在以有限生物医学数据集为特征的拉曼光谱应用中的机器学习鲁棒性。
英文摘要
Access to sufficiently large biomedical datasets remains a major obstacle for machine learning in Raman spectroscopy-based diagnostics. In particular, for glioma analysis, datasets are typically small and heterogeneous, affected by acquisition-specific variability. This work investigates the utility of deep generative augmentation in such a small-cohort setting. We analyze glioma biopsy spectra acquired from 58 tumor samples and consider both binary IDH-status classification and 6-class methylation subtype classification problems. To address the limited size and imbalance of the dataset, we develop a conditional variational autoencoder ($β$-CVAE) capable of generating class-conditioned synthetic Raman spectra. The generated data are evaluated in Train-on-Synthetic, Test-on-Real (TS/TR) and Train-on-Synthetic+Real, Test-on-Real (TSR/TR) settings under a strict patient-isolated cross-validation protocol. Models trained exclusively on synthetic data underperform models trained on real spectra, indicating a substantial domain gap between synthetic and real distributions. However, augmenting the real training data with synthetic spectra consistently improves classification performance across multiple models. These findings indicate that, even with a limited number of independent patient samples, generative models can capture sufficient structure to provide useful regularization for downstream classifiers. We also investigate a reconstruction-based inference strategy, termed Classification by Reconstruction (CbR), in which class prediction is based on reconstruction error under different class conditions. Overall, the results support the use of deep generative augmentation as a practical strategy for improving machine learning robustness in Raman spectroscopy applications characterized by limited biomedical datasets.