arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31635cs.LGcs.AI

下一事件准确率无法揭示的问题:急诊科轨迹模拟器的闭环评估

What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators

Zhen Xuen Brandon Low

AI总结:

针对临床轨迹模拟器,提出EDSim-Bench闭环评估框架,证明下一事件准确率不足以评估模拟性能,并揭示序列监督可显著降低发散度,强调多维度闭环评估的必要性。

AI中文摘要:

临床轨迹模型通常通过在观测历史数据上的下一事件准确率进行评估。然而,模拟任务有所不同:模型必须基于自身生成的事件进行条件化,这可能导致错误累积。尽管这一问题在序列建模领域广为人知,但尚未在临床轨迹模拟器中进行系统量化。我们开发了EDSim-Bench来评估这一失败模式,使用了MIMIC-IV-ED中的425,028次就诊记录,并在MC-MED上进行了外部复制验证,同时发布了评估协议和评分代码。从留出的就诊前缀开始,模型生成每次就诊的剩余部分,并在终止、事件组成、时间安排、条件保真度和占用率预测方面进行评估,以仅训练的三阶n-gram作为参考基线。尽管下一事件准确率差异在0.001以内,三种神经架构在展开(rollout)下的表现却截然不同。在不同随机种子下,一种Transformer配置的终止评分在0.43到0.96之间波动,其发散度是n-gram的4.2倍至137倍;没有一种前缀训练的神经模型在终止或事件组成方面接近n-gram的表现。推理时的干预措施改善了终止表现,但未能同时恢复事件组成和时间安排。对每个符合条件的序列位置进行监督,而非仅监督最终前缀位置,与Transformer、GRU和LSTM模型的发散度降低一至两个数量级相关,且这一模式在模型规模扩展、时间偏移和外部站点评估中持续存在。然而,即使是最佳模型,其生成的访问时长也约为观测时长的一半,并且模型排名在占用率预测(一项与床位管理相关的下游指标)上发生反转。这些结果表明,下一事件准确率不足以评估临床轨迹模拟器,并促使在多个随机种子、展开标准和下游任务上进行闭环评估。

英文摘要:

Clinical trajectory models are usually evaluated by next-event accuracy on observed histories. Simulation is different: models must condition on their own generated events, allowing errors to compound. Although this problem is well known in sequence modelling, it has not been systematically quantified for clinical trajectory simulators. We developed EDSim-Bench to evaluate this failure mode using 425,028 MIMIC-IV-ED stays, with external replication on MC-MED, and release the evaluation protocol and scoring code. Starting from held-out visit prefixes, models generate the remainder of each visit and are evaluated on termination, event composition, timing, conditional fidelity, and occupancy forecasting, with a train-only order-3 n-gram as a reference baseline. Despite next-event accuracies within 0.001, three neural architectures behaved very differently under rollout. Across seeds, one Transformer recipe ranged from 0.43 to 0.96 in termination score and from 4.2- to 137-fold the divergence of the n-gram; no prefix-trained neural model approached the n-gram on termination or event composition. Inference-time interventions improved termination but did not jointly recover composition and timing. Supervising every eligible sequence position rather than only the final prefix position was associated with one to two orders of magnitude lower divergence across Transformer, GRU, and LSTM models, with the pattern persisting under model scaling, temporal shift, and external-site evaluation. Nevertheless, even the best model generated visits approximately half as long as observed, and model rankings reversed on occupancy forecasting, a downstream quantity relevant to bed management. These results show that next-event accuracy is insufficient to evaluate clinical trajectory simulators and motivate closed-loop evaluation across seeds, rollout criteria, and downstream tasks.

↑