arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33117cs.LGcs.AIcs.CLeess.SP

ECG-Scroll:用于动态心电图解读的长时程、流式基准与智能体环境

ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms

Haitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li, Yiheng Mao, Zhengxing Huang

中文总结 AI 辅助

针对动态心电图长时程流式解读,提出 ECG-Scroll 基准与智能体环境,以在线决策方式评估记忆、工具使用和规划能力,并发布大规模数据集。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)现在能够解读标准的十秒、十二导联心电图(ECG),并具备临床依据和奖励验证的推理能力。然而,真实的 cardiac 监测并非如此。动态(Holter)和遥测记录持续数小时至数天,并且是边流入边读取的,其临床决定性发现是阵发性的、短暂的发作,埋藏在原本平淡无奇的轨迹中。这样的记录无法在诊断分辨率下完整放入一个上下文,而且其未来尚未发生,因此阅读者必须在线工作,决定现在测量什么,将证据在流逝时提交给记忆,并在事件发生时进行报告。我们将长时程心电图解读重新定义为一种长时程、在线(流式、因果)的序列决策过程,并引入 ECG-Scroll。作为基准,长的动态记录被分块流式传输给智能体,智能体必须定位、量化并及时标记阵发性事件,且无法访问未来信号;由于底层信号被保留,每个答案都可以对照客观真实值进行核查,从而提供基于规则而非基于评判者的奖励,并且流式公式增加了一个批量评估无法表达的指标,即事件发生与智能体记录该事件之间的检测延迟。作为智能体环境,它是一个固定的、gym 风格的交互层,锻炼了单次查看 ECG 模型从未触及的三项能力:记忆、通过基于信号的测量而非读取像素来使用工具,以及规划现在测量什么以及何时提交。我们发布了 390 个完整记录实例,涵盖 2,536 小时的双导联动态心电图,并在线评估了一个信号阈值规则智能体以及现成的 LLM 智能体,描述了它们如何使用记忆、工具和规划,以及基准的改进空间所在。

英文摘要

Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream in, and their clinically decisive findings are paroxysmal, brief episodes buried in an otherwise unremarkable trace. Such a recording cannot be held in one context at diagnostic resolution, and its future has not yet happened, so a reader must work online, deciding what to measure now, committing evidence to memory as it passes, and reporting events as they occur. We recast long-duration ECG interpretation as a long-horizon, online (streaming, causal) sequential decision process and introduce ECG-Scroll. As a benchmark, long ambulatory recordings are streamed to an agent chunk by chunk, and it must localize, quantify, and promptly flag paroxysmal events without access to future signal; because the underlying signal is retained, every answer is checkable against objective ground truth, giving rule-based rather than judge-based rewards, and the streaming formulation adds a metric batch evaluation cannot express, the detection latency between an event's onset and the moment the agent records it. As an agent environment, it is a fixed, gym-style interaction layer that exercises three competencies single-glance ECG models never touch: Memory, Tool use through signal-grounded measurement rather than reading pixels, and Planning of what to measure now and when to commit. We release 390 whole-recording instances spanning 2,536 hours of two-lead ambulatory ECG and evaluate a signal-threshold rule agent alongside off-the-shelf LLM agents online, characterizing how they use memory, tools, and planning and where the benchmark's head-room lies.

↑