AI 中文总结
研究针对大型睡眠研究库队列发现问题,提出基于逻辑的时间队列发现引擎,采用QEL和BEST,组织成三种查询模式,用Python实现原型并评估,2DFC构建索引性能优,队列选择查询高效,是符号生物医学计划一部分。
AI 中文摘要
大型睡眠研究存储库包含丰富的带时间戳的生理注释,但队列发现通常仍通过临时脚本或标量索引过滤器来实现。我们提出了一个基于逻辑的时间队列发现引擎,它将形式语义、模型检查、专门索引和实证评估引入到一个统一的生物医学信息学框架中。我们采用有理集合逻辑(QEL)作为睡眠数据查询的密集时间形式基础,并将每个带注释的多导睡眠图表示为生物医学事件结构时间模型(BEST),即从事件标签到不重叠有理区间集合的有限映射。队列发现被表述为对BEST数据库上的QEL公式进行模型检查。我们将常见的睡眠研究需求组织成三种可重复使用的时间查询模式:单事件检索、双事件时间模式匹配和事件数据提取。该原型队列发现引擎用Python实现,有内存和MongoDB支持的执行模式,并在包含多达9000万个区间的合成区间数据集以及来自克利夫兰儿童睡眠与健康研究(CCSHS)的包含515名受试者、202587个区间、23个事件标签的真实世界国家睡眠研究资源注释上进行了评估。2DFC在线性空间和线性构建时间内构建索引,将9000万个区间的构建时间从使用RTFC的11549秒和使用2DRT的23902秒减少到3655秒。在CCSHS上,队列选择查询在原生规模下以亚秒级延迟执行,在1000倍规模下在45秒内执行。这项工作是通讯作者倡导的符号生物医学计划的一部分。
英文摘要
Large sleep-study repositories contain rich time-stamped physiological annotations, but cohort discovery is still commonly implemented as ad hoc scripts or scalar-index filters. We present a logic-based temporal cohort discovery engine that brings formal semantics, model checking, specialized indexing, and empirical evaluation into a unified biomedical informatics framework. We adopt Rational Ensemble Logic (QEL) as a dense-time formal foundation for sleep-data querying and represent each annotated polysomnogram as a Biomedical Event Structure Temporal Model (BEST), a finite mapping from event labels to non-overlapping rational interval ensembles. Cohort discovery is formulated as model checking of QEL formulas over BEST databases. We organize common sleep-research requirements into three reusable temporal query patterns: single-event retrieval, dual-event temporal pattern matching, and event data extraction. The prototype cohort discovery engine was implemented in Python with in-memory and MongoDB-backed execution modes and evaluated on synthetic interval datasets containing up to 90 million intervals and on real-world National Sleep Research Resource annotations from the Cleveland Children's Sleep and Health Study (CCSHS) containing 515 subjects, 202,587 intervals, 23 event labels. 2DFC constructs indexes in linear space and linear build time, reducing build time at 90 million intervals from 11,549 s with RTFC and 23,902 seconds with 2DRT to 3,655 seconds. On CCSHS, cohort-selection queries executed at sub-second latency at native scale and under 45 seconds at 1,000 times scale. This work is a part of the Symbolic Biomedicine program championed by the corresponding author.
CommentsUnder review by the Journal of Biomedical Informatics (JBI)