arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ECGQuest:用于心电图的语言模型基准测试与微调

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni

arXiv 2608.30893首次发表:更新:

发表机构

Emory University; Georgia Institute of Technology(埃默里大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究开发了ECGQuest基准数据集,对语言模型进行ECG知识的零样本评估与微调,发现参数高效微调可让小模型与大商业模型竞争,通用模型在ECG任务中优于医学专用模型。

AI 中文摘要

心电图(ECG)解读需要心脏病学、电生理学、临床诊断、ECG波形、信号采集及仪器等多领域知识。然而现有的语言模型基准主要评估广泛的医学知识或单个ECG信号与图像的解读,而非ECG解读所需的更广泛背景知识。我们开发了ECGQuest,这是一个基于文献的资源,用于评估和微调ECG专用语言模型。基于GPT-4o的管道从23篇ECG文献及2003-2025年的《Computing in Cardiology》会议论文中生成问题,最终数据集包含10904个独特的判断题及其否定形式(共21808个问答对)。我们在保留的测试集上以零样本设置评估了3个商业语言模型和20个开源语言模型,对5个参数规模为7-14B的开源模型使用Low-Rank Adaptation进行微调,同时纳入BERT和BiomedBERT作为监督编码器基线。我们在MedMCQA和MedQA的ECG相关子集(已通过官方答案键转换为判断题)上评估泛化能力。ECGQuest上的零样本准确率范围为49.5%至74.4%,GPT-5表现最佳;通用模型优于医学专用模型,多个模型表现出强烈的判断题偏向,编码器基线表现接近随机水平。微调使所有开源模型的准确率提升了6.5-14.1%,微调后的DeepSeek-R1-Distill-Qwen-14B达到76.3%的准确率,而5个模型的投票集成达到78.5%。在MedMCQA和MedQA上,微调主要使较弱或存在类别偏向的模型受益,并未持续提升强基础模型的性能。ECGQuest提供了一个可复现的ECG背景知识基准,且参数高效微调可使较小的语言模型与规模大得多的商业模型竞争。

英文摘要

Electrocardiogram (ECG) interpretation requires knowledge of cardiology, electrophysiology, clinical diagnosis, ECG waveforms, signal acquisition, and instrumentation. Existing language-model benchmarks, however, primarily assess broad medical knowledge or interpretation of individual ECG signals and images rather than the broader contextual knowledge required for ECG interpretation. We developed ECGQuest, a literature-grounded resource for evaluating and fine-tuning ECG-specific language models. A GPT-4o-based pipeline generated questions from 23 ECG references and Computing in Cardiology proceedings from 2003-2025. The final dataset contains 10,904 unique True/False questions paired with their negated forms (21,808 Q&A pairs). We evaluated three commercial and 20 open-source language models on a held-out test set in a zero-shot setting. Five open-source models with 7-14B parameters were fine-tuned using Low-Rank Adaptation, with BERT and BiomedBERT included as supervised encoder baselines. Generalization was assessed on ECG-related subsets of MedMCQA and MedQA converted to binary True/False questions using official answer keys. Zero-shot accuracy on ECGQuest ranged from 49.5% to 74.4%, with GPT-5 performing best. General-purpose models outperformed medically specialized models, several models showed strong True/False bias, and encoder baselines performed near chance. Fine-tuning improved all open-source models by 6.5-14.1%. Fine-tuned DeepSeek-R1-Distill-Qwen-14B reached 76.3% accuracy, while a five-model voting ensemble reached 78.5%. On MedMCQA and MedQA, fine-tuning mainly benefited weaker or class-biased models and did not consistently improve strong base models. ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑