arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMTClinic:面向临床领域的多模态、多语言时间序列问答与推理基准

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh

arXiv 2609.04842首次发表:更新:

发表机构

Institute of Engineering and Management; IIT Patna; University of Calcutta; Bangabasi Morning College(工程与管理学院; 巴特那印度理工学院; 加尔各答大学; 邦加巴西晨学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向临床领域的多模态多语言时间序列问答与推理基准MMTClinic,融合多类数据、覆盖五种语言与三项临床任务,评估13种LLMs并揭示其性能差异,为相关医疗AI研究提供宝贵资源。

AI 中文摘要

临床场景中的时间序列数据对捕捉患者健康随时间的动态变化、实现及时诊断、个性化治疗及早期危重事件检测至关重要。然而,开发临床可靠且语言包容的医疗AI系统仍是重大挑战,核心原因是缺乏能反映真实临床场景复杂性的多模态、多语言且以时间序列为基础的基准。为填补这一空白,本文提出MMTClinic,一个用于评估大语言模型(LLMs)在涉及临床时间序列的复杂推理与问答任务上的基准。MMTClinic融合文本、医学图像及多变量生理信号,包含3万条问答对(1.5万条多项选择题(MCQs)和1.5万条开放式问题),覆盖英语、印地语、孟加拉语、马拉地语、泰米尔语五种语言。这些问题涵盖三项重要临床任务——死亡率预测、心率预测及SOFA评分估算。我们在零样本、少样本及思维链设置下评估了13种最先进的LLMs,评估显示模型在不同任务、语言及模态间存在显著性能差异,凸显了当前临床推理能力的局限性。MMTClinic为推进多语言、多模态及时间序列感知的医疗AI研究提供了宝贵资源,数据集将在本研究成功接收后公开。

英文摘要

Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑