arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用模糊认知图的可解释合成医学表格数据生成,用于临床决策支持

Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps

Michael Vasilakakis, Dimitris K. Iakovidis

arXiv 2610.00391首次发表:更新:

发表机构

University of Thessaly(色萨利大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于模糊认知图的合成医学表格数据生成框架,通过可解释模糊集和因果边权重保留临床依赖,在TSTR协议下达到高准确率和隐私保护,为临床决策支持提供透明高效方案。

AI 中文摘要

合成医学表格数据的生成对于开发和验证基于计算机的医学系统(CBMSs)至关重要,尤其是在真实临床数据因隐私、伦理或数据可用性限制而受到约束的情况下。现有的概率和深度生成模型往往缺乏可解释性,并且无法保留具有临床意义的依赖关系,限制了它们在安全关键应用中的适用性。本文提出了一种新颖的模糊认知图(FCMs)应用框架,用于合成医学表格数据生成,具有明确的因果性和隐私保护。临床特征使用语言上可解释的模糊集进行描述,特征间依赖关系被编码为从模糊集交集计算的FCM边权重。合成患者记录通过将随机初始化的语言激活向量传播通过FCM直至收敛,然后进行去模糊化以生成临床上连贯的数值。该方法天然处理混合数据类型和健康记录中常见的领域约束。在UCI医学基准数据集上的实验评估表明,在训练于合成测试于真实(TSTR)协议下,该方法表现出竞争性性能。所提出的方法在心脏病数据集上实现了高达0.81的准确率和高达0.90的AUROC,匹配或超过TVAE和高斯Copula基线,同时仅使用CPU运行。保真度指标包括KS互补(高达0.91)和相关性相似度(高达0.95)证实了强统计连贯性,DCR基线保护分数始终超过TVAE,确认了充分的隐私保证。这些结果表明,基于因果的可解释模糊建模为CBMSs中可信合成数据生成提供了一种计算高效且透明的替代深度生成模型的方案。

英文摘要

Synthetic medical tabular data generation has become essential for developing and validating computer-based medical systems (CBMSs) when real clinical data is restricted due to privacy, ethical, or data availability limitations. Existing probabilistic and deep generative models often lack interpretability and fail to preserve clinically meaningful dependencies, limiting their suitability for safety-critical applications. This paper proposes a novel application of Fuzzy Cognitive Maps (FCMs) in a framework for synthetic medical tabular data generation with explicit causality and privacy preservation. Clinical features are described using linguistically interpretable fuzzy sets, and inter-feature dependencies are encoded as FCM edge weights computed from fuzzy set intersections. Synthetic patient records are generated by propagating randomly initialized linguistic activation vectors through the FCM until convergence, followed by defuzzification to produce clinically coherent numerical values. The approach natively handles mixed data types, and domain constraints common in health records. Experimental evaluation on UCI medical benchmark datasets demonstrates competitive performance under a Train-on-Synthetic-Test-on-Real (TSTR) protocol. The proposed method achieves accuracy of up to 0.81 and AUROC of up to 0.90 on the Heart Disease dataset, matching or exceeding TVAE and Gaussian Copula baselines while running exclusively on CPU. Fidelity metrics including KS Complement (up to 0.91) and Correlation Similarity (up to 0.95) confirm strong statistical coherence, and DCR Baseline Protection scores consistently exceed those of TVAE, confirming adequate privacy guarantees. These results demonstrate that causally grounded, interpretable fuzzy modeling offers a computationally efficient and transparent alternative to deep generative models for trustworthy synthetic data generation in CBMSs.

Comments6 pages, 2 figures, 1 table. Accepted for publication in the 2026 IEEE 39th International Symposium on Computer-Based Medical Systems (CBMS), Limassol, Cyprus

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑