生物医学知识构建:软件工程视角
Biomedical Knowledge Composition: A Software Engineering Perspective
浏览论文内容
中文总结 AI 辅助
本文从软件工程视角,阐述生物医学知识图谱的核心地位与相关挑战,提出需借鉴Web工程的可组合可复现实践,为生物医学知识基础设施构建提供研究议程。
中文摘要 AI 辅助
生物医学研究已积累了海量分子、临床及人群数据,但将这些数据转化为可操作知识仍受技术与组织层面的难题制约。本文对生物医学知识基础设施的两种视角进行了统一处理。第一种视角面向软件工程师介绍生物医学领域:解释为何知识图谱(KGs)是现代生物医学中核心的整合数据结构,阐述了五项数据协调挑战(标识符映射、实体解析、模式对齐、证据整合及溯源追踪),梳理了从药物发现到数字孪生的应用领域,并介绍了六个具有不同选择的代表性知识图谱系统。第二种视角探讨了构建生物医学知识基础设施为何仍如此困难。我们认为一个重要的根本原因是,生物医学领域对能使其他成熟领域(尤其是Web工程)开发具备可靠可组合性与可复现性的软件工程工具与实践的采用有限,这些工具与实践包括包管理、类型化命名空间、规范交换格式、服务组合协议、可复现管道及生命周期治理。在此背景下,本文列举了生物医学数据整合的八项开放性工程挑战,每项挑战都有部分解决方案,但尚未形成普遍采用的技术栈。关键的是,本文将重点从描述已部署的知识图谱实例转向构建这些实例的可复现过程:可复用的构建管道、版本化依赖项,以及让他人能够从源代码编译和定制知识图谱而非仅使用静态制品的工程实践。两种视角共同为新手提供了领域基础,并为希望对生物医学知识基础设施做出变革性贡献的软件工程师提供了研究议程。
英文摘要
Biomedical research has accumulated vast molecular, clinical, and population data, yet translating this wealth into actionable knowledge remains constrained by technical and organizational difficulties. This article presents a unified treatment of two perspectives on biomedical knowledge infrastructure. The first introduces the biomedical domain to software engineers: it explains why knowledge graphs (KGs) are the central integrative data structure in modern biomedicine, characterizes five data harmonization challenges (identifier mapping, entity resolution, schema alignment, evidence integration, and provenance tracking), surveys application domains from drug discovery to digital twins, and profiles six representative KG systems with contrasting choices. The second perspective asks why engineering biomedical knowledge infrastructure remains so difficult. We argue that a contributing root cause is limited adoption of software tooling and practices that make development in other mature domains - particularly web engineering - reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. Against this backdrop, eight open engineering challenges for biomedical data integration are catalogued, each with partial solutions but no universally adopted stack. Crucially, the article shifts emphasis from describing deployed KG instances toward the reproducible process of assembling them: reusable build pipelines, versioned dependencies, and engineering practices that let others compile and customize a KG from source rather than consuming a static artifact. Together, the two perspectives provide domain grounding for newcomers and a research agenda for software engineers seeking to make transformative contributions to biomedical knowledge infrastructure.