Holtercare-Bench:用于评估长期动态心电图分析的多模态基准
Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
浏览论文内容
中文总结 AI 辅助
该研究针对现有多模态大语言模型在长期动态心电图分析中的不足,构建了含22980个问答对的Holtercare-23K数据集及Holtercare-Bench基准,评估发现零样本模型处理超长病理序列性能差,微调后提升显著,为长期医疗多模态大语言模型提供基础基准。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLMs)在医疗应用中表现出色,但大多数模型更倾向于静态图像或短期信号。在关键的动态心电图(ECG)领域,由于缺乏高质量数据集和基准,模型在复杂的时间推理和诊断报告生成方面存在困难。为解决这一问题,我们推出了(i)Holtercare-23K,这是一个大规模多模态动态ECG数据集,包含来自788条临床动态心电图(Holter)记录的22980个问答对,具有新颖的信号-视频-文本三模态对齐特性。基于该数据集,我们提出了(ii)Holtercare-Bench,这是一个多模态基准,用于评估模型的时间定位、临床诊断和全局总结能力。对领先MLLMs的零样本评估显示,其在处理超长病理序列时存在显著性能差距,但对代表性模型进行微调后可获得大幅提升。本研究阐明了当前MLLMs在电生理学领域的局限性,并为长期医疗MLLMs提供了基础基准,项目地址为该URL。
英文摘要
While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks. To address this, we introduce (i) Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal-video-text tri-modal alignment. Based on this dataset, we present (ii) Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long-term medical MLLMs. Our project is available at https://github.com/ZJU4HealthCare/Holtercare-Bench.
发表机构
- Zhejiang University(浙江大学)
- Beijing Institute of Technology(北京理工大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。