arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于五维多表分析的方法与概念框架:复杂数据复用的统一方法

Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

Edouard Lansiaux, Hugo Kazzi, Aurélien Loison, Slim Hammadi, Emmanuel Chazard

arXiv 2608.26149首次发表:更新:

发表机构

CHU de Lille; Lille Centrale Institute; Lille University(里尔大学中心医院; 里尔中央理工学院; 里尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出关系超图Transformer(RHT)架构,用于解决复杂关系数据的多表学习问题,在Synthea数据集上验证其生成语义连贯嵌入的能力,计划后续在MIMIC-IV上开展临床验证。

AI 中文摘要

多表学习仍是医疗及其他复杂信息系统机器学习领域的重大挑战。关系数据兼具多类复杂性,包括大数据量、高维变量、高基数分类特征、复杂表间依赖关系及重复时间观测值。本文提出关系超图Transformer(Relational Hypergraph Transformer, RHT),该统一架构将关系数据库表示为超图,学习五维嵌入(PentE),并执行稀疏关系注意力,其复杂度与平均关系度成正比,而非实体数量的平方。本文对该架构进行形式化定义,推导其注意力机制的复杂度,并提供开源参考实现。我们在公开的Synthea合成电子健康记录数据集上评估RHT,任务为按就诊预测SNOMED CT病症代码的多标签分类,该任务具有高分类基数和长尾标签分布特征。与表格、关系型及时间图基线的对比显示,RHT生成的嵌入语义更连贯,同时保持计算可扩展性。在该基准中,XGBoost实现了最高的稀有代码召回率,而RHT则获得最强的嵌入语义连贯性。我们还通过消融研究量化了各架构组件的贡献。在获得PhysioNet认证后,计划对MIMIC-IV开展临床验证,源代码和实验方案已随附在相关代码库中。

英文摘要

Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features, complex inter-table dependencies, and repeated temporal observations. We introduce the Relational Hypergraph Transformer (RHT), a unified architecture that represents relational databases as hypergraphs, learns pentadimensional embeddings (PentE), and performs sparse relational attention with complexity proportional to the average relational degree rather than the square of the number of entities. We formally define the architecture, derive the complexity of its attention mechanism, and provide an open-source reference implementation. We evaluate RHT on the public Synthea synthetic electronic health record dataset using multi-label prediction of SNOMED CT condition codes per encounter, a task characterized by high categorical cardinality and long-tailed label distributions. Comparisons with tabular, relational, and temporal graph baselines show that RHT produces more semantically coherent embeddings while remaining computationally scalable. In this benchmark, the highest rare-code recall is achieved by XGBoost, whereas RHT attains the strongest embedding semantic coherence. We also report ablation studies quantifying the contribution of each architectural component. Clinical validation on MIMIC-IV is planned following PhysioNet credentialing. Source code and experimental protocols are provided in the accompanying repository.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑