arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

失败揭示指标遗漏之处:用于心电图分类器递归优化的证据驱动智能体

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

Jinliang Deng, Yiming Niu, Yibo Pan, Zhiqi Shao, Qin Luo, Yongxin Tong

arXiv 2607.24419首次发表:更新:

发表机构

Beihang University; Chongqing University; Fuwai Hospital(北京航空航天大学; 重庆大学; 阜外医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究聚焦心电图分类器优化依赖人工的问题,提出RecursiveECG框架,利用LLM基于具体失败情况和心电图证据优化分类器,经特殊转换和审查分析制定修订,在多数据集上表现出色,验证了基于证据优化过程的有效性。

AI 中文摘要

深度模型推动了12导联心电图分类的发展,但其优化仍严重依赖人类专家检查失败情况并迭代修订分类器设计。基于语言模型(LLM)的智能体虽有自动化模型设计潜力,但仅靠总体性能指标缺乏对个体失败原因及分类器修订方式的洞察。我们提出RecursiveECG,这是一个证据驱动的以LLM为设计者的框架,LLM作为离线模型设计者,基于具体失败情况和客观心电图证据优化心电图分类器。通过标准到测量编译将精选的心电图标准转换为经过验证的确定性函数,为个体心电图生成可重复、有参考依据的测量值。在此基础上,基于证据的失败审查通过联合考虑原始波形、测量值和模型输出分析失败及对比案例,使LLM能够诊断分类器局限性并制定有针对性的修订。候选修订在固定问题契约下执行并重新评估,仅保留有证据支持的更新。在PTB-XL、Georgia和CPSC2018数据集上,RecursiveECG始终优于强大的基线,平均相对提升10.0%。广泛的消融和迁移研究进一步验证了其基于证据的优化过程的有效性。

英文摘要

Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier should be revised. We present RecursiveECG, an evidence-driven LLM-as-Designer framework in which an LLM serves as an offline model designer that refines ECG classifiers based on concrete failures and objective ECG evidence. To ground failure diagnosis in executable evidence, Criteria-to-Measurement Compilation converts curated ECG criteria into validated deterministic functions that produce reproducible, reference-backed measurements for individual ECGs. Building on these measurements, Evidence-Grounded Failure Review analyzes failed and comparator cases by jointly considering raw waveforms, measurements, and model outputs, enabling the LLM to diagnose classifier limitations and formulate targeted revisions. Candidate revisions are executed and re-evaluated under a fixed problem contract, and only evidence-supported updates are retained. The resulting predictor is frozen after refinement and requires no LLM inference during deployment, while an audit trail links each accepted revision to its supporting evidence. Across PTB-XL, Georgia, and CPSC2018, RecursiveECG consistently outperforms strong baselines, achieving an average relative improvement of 10.0%. Extensive ablation and transfer studies further validate the effectiveness of its evidence-grounded refinement process.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑