发表机构
Old Dominion University; North Carolina Agricultural and Technical State University(奥多明尼昂大学; 北卡罗来纳农业技术州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生物医学数据中标签缺失和复杂依赖问题,提出CopDAG框架,结合copula模型与DAG因果发现,在16个数据集上聚类性能领先。
AI 中文摘要
诊断错误和错误标注在生物医学中很常见,这损害了预测模型和数据驱动结果的可靠性。基于特征之间的复杂关系对未标注的生物医学数据进行分层,消除了对数据标签的需求,并克服了监督学习的局限性。传统的聚类方法假设了限制性的数据分布,使其在捕获高维生物医学数据中的复杂依赖关系时表现欠佳。本文介绍了一种新颖的、对聚类友好的数据表示框架,该框架将 copula 模型的非高斯和非线性特征依赖性与基于有向无环图(DAG)的因果结构发现(CSD)方法的集成相结合。copula 模型通过放宽与多元正态性、线性依赖和对称关系相关的假设来建模灵活的多元分布,而基于 DAG 的 CSD 方法集成则识别特征之间稳定的因果关系。当使用 K-means 聚类时,通过所提出的 copula 自适应 DAG(CopDAG)获得的新数据表示在 16 个生物医学数据集上的归一化聚类准确率和调整兰德指数方面,在 12 种方法中排名第一。我们的 CopDAG 方法直接根据特征关系预测真实类别标签,无需数据标注和监督学习,同时还提供聚类可视化和生物医学数据特征的可解释因果结构。
英文摘要
Diagnostic errors and mislabeling are common in biomedicine, which compromise the reliability of predictive models and data-driven outcomes. Stratifying unlabeled biomedical data based on complex relationships between features eliminates the need for data labels and overcomes the limitations of supervised learning. Traditional clustering methods assume restrictive data distributions, making them suboptimal for capturing complex dependencies in high-dimensional biomedical data. This paper introduces a novel cluster-friendly data presentation framework that integrates the non-Gaussian and non-linear feature dependence of copula models with an ensemble of causal structure discovery (CSD) methods based on Directed Acyclic Graphs (DAGs). While copulas model flexible multivariate distributions by relaxing assumptions related to multivariate normality, linear dependence, and symmetric relationships, an ensemble of DAG-based CSD methods identifies stable causal relationships between features. When clustered using K-means, the new data representation obtained by the proposed copula-adapted DAG (CopDAG) ranks first among the 12 methods in normalized clustering accuracy and adjusted Rand index across 16 biomedical datasets. Our CopDAG method predicts ground-truth class labels directly from feature relationships without data annotations and supervised learning, while also providing cluster visualizations and explainable causal structures of the biomedical data features.