arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32740cs.CV

AnesTRACE:从多模态感知到多步决策的术中麻醉基准测试

AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making

Ziwei Huang, Qi Gao, Zhe Ji, Yuanyuan Yao, Fengjiang Zhang, Min Yan, Zhongle Xie, Gang Chen

AI总结:

本文提出AnesTRACE基准套件,评估术中麻醉从多模态感知到多步决策的能力,发现现有模型在细粒度视觉定位和干预选择上仍存在困难,强调安全与及时性的重要性。

AI中文摘要:

术中麻醉要求系统能够解读不断演变的多模态证据,推荐及时的管理措施,并根据患者状态的变化修订决策,然而现有基准通常孤立地评估感知或单点推理。我们引入了AnesTRACE,一个包含AnesTRACE-Bench和AnesTRACE-Eval的评估套件。基于公开的围手术期数据集并辅以麻醉医师标注,AnesTRACE-Bench评估术中感知、单点麻醉决策和多步麻醉决策。AnesTRACE-Eval通过麻醉医师定义的临床正确性、证据依据、任务完整性和安全性标准评估开放式响应,并对多步决策评估时间一致性;其领域特定评估器通过监督微调和基于专家评审判断的偏好对齐进行训练。在超过30个模型中,细粒度视觉定位和干预选择仍然困难:领先模型在TEE视觉定位上仅达到32.2 mIoU,并在多步管理中保持17.5%的重大/关键安全错误率。评估器与麻醉医师的一致性在两个训练阶段均有所提高,而最佳决策质量伴随74.3秒的P95延迟。这些结果表明,仅凭总体性能并不能确立安全、及时的纵向决策。我们在以下https URL发布代码。

英文摘要:

Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesTRACE-Bench and AnesTRACE-Eval. Built from public perioperative datasets with anesthesiologist annotation, AnesTRACE-Bench evaluates Intraoperative Perception, Single-point Anesthesia Decision-Making, and Multi-step Anesthesia Decision-Making. AnesTRACE-Eval assesses open-ended responses through anesthesiologist-defined criteria for Clinical Correctness, Evidence Grounding, Task Completeness, and Safety, with Temporal Consistency for multi-step decisions; its domain-specific evaluator is trained by supervised fine-tuning and preference alignment on expert-reviewed judgments. Across more than 30 models, fine-grained visual grounding and intervention selection remain difficult: the leading model reaches only 32.2 mIoU for TEE visual grounding and retains a 17.5\% Major/Critical Safety Error Rate in multi-step management. Evaluator alignment with anesthesiologists improves across both training stages, while the best decision quality is accompanied by a 74.3-second P95 Latency. These results show that aggregate performance alone does not establish safe, timely longitudinal decision-making. We release our code at https://zjudbxai.github.io/AnesTRACE/.

↑