发表机构
Roshan AI; University of Arizona(Roshan AI; 亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将ICU警报削减重构为三向分诊,提出认证式方法在95%置信度下控制真实警报静音比例,在VTaC基准上抑制74.8%虚假警报且仅静音1.5%真实警报,并揭示多重性校正对认证结果的代价。
AI 中文摘要
在VTaC基准中,71%的室性心动过速警报是虚假的,但静音一个真实警报可能会延迟对危险心律失常的识别。我们将警报削减重新定义为三向分诊(保留、抑制或延迟),并界定此分析视为有害的决策界限:在抑制的警报中,真实警报的比例在独立同分布事件采样下,以95%的置信度保持在用户设定的预算以下。共享同一波形记录的警报是相互依赖的,因此聚类分析作为敏感性检查。在官方划分上,5%的预算在所有三个种子中均获得认证,抑制了74.8%的虚假警报,同时静音了1.5%的真实警报,AUROC为0.953,挑战得分为83.33,数值上与已发表的十一个系统中最强者相当。我们的核心发现衡量了多重性带来的代价:校正对每个候选都收费,因此更细的网格可能认证更少。在留出校准下,我们声明的885单元网格在15次折叠运行中认证了1次,而在单独的选择划分上选择网格则认证了8次。我们预测每个预算所需的校准量,使不可认证的预算成为设计参数。最后,向策略网格添加学习到的可靠性维度并未使认证边界更尖锐。
英文摘要
In the VTaC benchmark 71% of ventricular-tachycardia alarms are false, but silencing a real one can delay recognition of a dangerous arrhythmia. We reframe alarm reduction as three-way triage (retain, suppress, or defer) and bound the decision this analysis treats as harmful: among suppressed alarms, the fraction that were genuine stays below a user-set budget with 95% confidence, under i.i.d. event sampling. Alarms sharing a waveform record are dependent, so the clustered analysis is a sensitivity check. On the official split a 5% budget certifies in all three seeds, suppressing 74.8% of false alarms while silencing 1.5% of genuine ones, at AUROC 0.953 and Challenge Score 83.33, numerically comparable to the strongest of the eleven published systems. Our central finding measures what multiplicity costs: the correction charges for every candidate, so a finer grid can certify strictly less. Under held-out calibration the 885-cell grid we declared certifies 1 of 15 fold-runs, while choosing the grid on a separate selection partition certifies 8. We project the calibration volume each budget needs, making an uncertifiable budget a design parameter. Finally, adding a learned reliability dimension to the policy grid did not sharpen the certified frontier.