CLIR-Bench:针对不规则临床时间序列的多模态问答基准测试
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
浏览论文内容
中文总结 AI 辅助
研究针对临床时间序列稀疏、不规则等问题,引入CLIR-Bench基准,通过四阶段流程构建,含6600个问答实例。实验发现现有通用模型处理稀疏临床证据困难,强调需更强不规则时间序列推理方法。
中文摘要 AI 辅助
临床时间序列对于患者监测、风险评估和临床决策支持至关重要。但它们往往稀疏、采样不规则且异步,模型难以识别临床问答所需的时间证据。现有基准主要关注规则采样时间序列问答或静态数据的医学问答,很少评估模型能否在不规则时间观测中可靠地得出答案。为填补这一空白,我们引入CLIR-Bench,这是一个通过有原则的四阶段流程从去标识的ICU记录构建的不规则临床时间序列问答基准。它包含6600个跨11个临床变量的问答实例,分为四个能力维度和11个任务。每个问题都与明确的时间证据和特定任务的答案推导规则相关联,可评估答案准确性和证据使用情况。实验表明现有通用模型在检索和推理稀疏临床证据方面存在困难,凸显了更强的不规则时间序列推理方法的必要性。我们的代码和数据可通过此https链接获取。
英文摘要
Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA). Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. CLIR-Bench contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks. Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use. Experiments show that existing generalist models struggle to retrieve and reason over sparse clinical evidence, highlighting the need for stronger irregular time-series reasoning methods. Our code and data are available at https://huggingface.co/datasets/winall/CLIR-Bench.
发表机构
- Shandong University(山东大学)
- University of Auckland(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。