发表机构
Cancerscan Inc.; Kyoto University; Hiroshima University(Cancerscan公司; 京都大学; 广岛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出统一预训练框架,整合层次子词聚合、部分掩码和交叉引用机制,提升医学表示学习,并在药物重定位中实现数据驱动的假设生成与排序。
AI 中文摘要
从电子健康记录和医疗索赔数据中的医学代码序列进行表示学习已在多种临床应用中取得成功,例如疾病预测。然而,将这种方法扩展到科学假设发现仍面临重大挑战。原因之一是许多现有的基于BERT的模型未能充分捕捉医学代码的层次结构以及诊断与治疗之间的复杂交互。为解决这些局限性,我们提出了一种新的统一预训练框架,该框架明确整合了层次子词聚合、部分掩码和交叉引用机制。所提出的模型在预训练目标和下游临床事件预测任务(包括痴呆症发病和住院)上均持续优于现有方法。我们还开展了一项针对阿尔茨海默病的计算机模拟药物重定位案例研究。在假设生成步骤中,我们的方法以数据驱动的方式成功重新发现了已知的有前景药物,而无需依赖文献等外部知识来源。随后,在假设优先级排序步骤中,我们引入了一种任务自适应表示方法,以缓解诊断向量中对历史处方信息的过度编码,从而实现对生成假设的稳健排序。本研究建立了一个基于观察性关联的假设生成与排序的探索性筛选工作流。重要的是,该框架并非旨在提供因果证据,而是识别有前景的候选药物以供后续严格的因果推断。总体而言,本研究表明,结合领域信息的表示学习与任务自适应表示控制能够实现实用的假设发现工作流。
英文摘要
Representation learning from medical code sequences in electronic health records and medical claims data has been successful in various clinical applications, such as those regarding disease prediction. However, significant challenges remain in extending this approach to the discovery of scientific hypotheses. One reason is that many existing BERT-based models fail to adequately capture the hierarchical structure of medical codes and the complex interactions between diagnoses and treatments. To address these limitations, we propose a new unified pre-training framework that explicitly integrates hierarchical sub-token aggregation, partial masking, and cross-reference mechanisms. The proposed model consistently outperformed existing methods on both pre-training objectives and downstream clinical event prediction tasks, including the onset of dementia and hospitalization. We also conducted an in silico drug repositioning case study targeting Alzheimer's disease. In the hypothesis generation step, our approach successfully rediscovered known promising drugs in a data-driven manner without relying on such external knowledge sources as the literature. Subsequently, in the hypothesis prioritization step, we introduced a Task-Adaptive Representation Approach to alleviate the over-encoding of historical prescription information within diagnostic vectors, enabling the robust prioritization of generated hypotheses. This study establishes an exploratory screening workflow for hypothesis generation and prioritization based on observational associations. Importantly, this framework is not intended to provide causal evidence, but rather to identify promising candidates for subsequent rigorous causal inference. Overall, this study demonstrates that domain-informed representation learning combined with task-adaptive representation control can enable a practical hypothesis discovery workflow.
CommentsAccepted at ICML 2026 AI for Science Workshop