发表机构
University Hospital Frankfurt, Goethe University Frankfurt; Goethe University Frankfurt; The Hessian Center for Artificial Intelligence (hessian.AI)(法兰克福大学医院,法兰克福大学; 法兰克福大学; 黑森人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究开发了问题驱动框架,通过PubMed API检索文献,用大语言模型提取队列名称,在青少年攻击行为遗传学用例中,补充了队列目录未覆盖的17个合格队列。
AI 中文摘要
背景:规划多研究分析需要识别具有相关参与者、表型和数据模态的队列,该过程通常依赖先验知识、队列目录和手动文献检索。我们开发了一种互补的问题驱动框架,用于检索相关科学文献并提取明确的队列名称。方法:该框架首先从可配置词汇表和模板生成多个PubMed查询,通过PubMed API自动检索得到的科学文献;随后,大语言模型使用针对研究问题定制的提示筛选检索到的标题和摘要,并提取明确的队列名称,提取的名称会进行去重并经人工审核。可配置代码、提示和示例输出可在指定URL获取。评估:作为用例,我们将该框架应用于青少年攻击行为遗传学研究,从5400个生成的PubMed查询中,框架检索到5254条唯一记录并识别出188个候选队列;使用预定义标准(包括参与者年龄和遗传数据可用性)进行人工筛选后,保留了44个合格队列;基于大语言模型的自动名称提取与人工标注者的一致性处于可接受范围。我们还使用相同研究问题检索了四个已建立的队列目录,其合并结果包含44个合格队列中的27个,另有17个未被任何队列目录检索到。结论:该框架将特定研究问题的词汇通过大规模自动文献检索转换为可筛选的队列清单,可跨人群、表型、数据模态和研究设计进行适配,为人工整理的队列目录提供基于文献的补充。
英文摘要
Background: Planning multi-study analyses requires identifying cohorts with the relevant participants, phenotypes, and data modalities. This process commonly relies on prior knowledge, cohort catalogues, and manual literature searches. We developed a complementary question-driven framework that searches relevant scientific literature and extracts explicit cohort names. Methods: The framework first generates multiple PubMed queries from configurable vocabularies and templates and retrieves the resulting scientific literature automatically through the PubMed API. A large language model then screens the retrieved titles and abstracts and extracts explicit cohort names using a prompt tailored to the research question. The extracted names are deduplicated with human review. Configurable code, prompts, and example outputs are available at https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery. Evaluation: As a use case, we applied the framework to youth aggression genetics. From 5,400 generated PubMed queries, the framework retrieved 5,254 unique records and identified 188 candidate cohorts. Manual screening using predefined criteria, including participant age and genetic-data availability, retained 44 eligible cohorts. Automated LLM-based name extraction was within the agreement range of human annotators. We also searched four established cohort catalogues using the same research question. Their combined results contained 27 of the 44 eligible cohorts, while 17 were not returned by any cohort catalogue search. Conclusion: The framework converts research-question-specific vocabulary into screenable cohort inventories via a large, automated literature search. It can be adapted across populations, phenotypes, data modalities, and study designs, and provides a literature-based complement to curated cohort catalogues.