arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedImageOSWorld:医学影像控制台GUI代理基准测试

MedImageOSWorld: Benchmarking GUI Agents for Medical Image Consoles

Ziyang Long, Xinqi Li, Lujing Xing, Hsin-Jung Yang

arXiv 2610.04800首次发表:更新:

发表机构

Biomedical Imaging Research Institute, Cedars-Sinai Medical Center; University of California, Los Angeles; Technical University of Munich; Berlin Ultrahigh Field Facility, Max Delbrück Center for Molecular Medicine in the Helmholtz Association(西达赛奈医疗中心生物医学成像研究所; 加利福尼亚大学洛杉矶分校; 慕尼黑工业大学; 亥姆霍兹协会马克斯·德尔布吕克分子医学中心柏林超高场设施)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MedImageOSWorld基准,在模拟CT/MR/超声控制台中评估GUI代理的医学采集能力,发现开放权重代理工作流完成度高但采集成功率低,感知与规划间存在显著性能鸿沟。

AI 中文摘要

图形控制台为医学采集辅助提供了实用的界面,使代理能够通过人类操作员所使用的控制和视觉反馈来工作。可靠的辅助需要将屏幕上的解剖结构与采集决策联系起来,这些决策决定了接下来可获得哪些影像证据。我们引入了MedImageOSWorld,一个用于在模拟CT、MR和超声控制台中评估此能力的基准。代理使用屏幕截图以及鼠标和键盘操作,配置协议、规划采集、检查生成的图像,并在七个能力级别(从控制台操作到反馈驱动控制)中进行纠正性调整。评估将特定于任务的工作流检查与隐藏的解剖学地面真值相结合,以分别评估程序完成情况和采集结果。一个通用的评估协议规定了情节条件和交互预算,而记录的轨迹支持分析代理如何观察、行动以及对采集反馈做出响应。在十一个开放权重代理中,成功率在0-100的范围内从3.0到25.0不等,而工作流进展率则达到21.5-74.5:代理完成了大部分控制台工作流,但很少采集到预期的解剖结构。成功在感知和规划之间崩溃,从最佳代理在最低三个级别上的79-90%,到毫米级规划最多12%,以及开放权重代理在闭环控制上的0%。两个专有代理的成功率约为34,并且主要在采集质量上超过最佳开放权重代理(61对42)。MedImageOSWorld提供了一个受控环境,用于研究通用GUI代理是否能够将视觉观察转化为有效的医学采集决策。

英文摘要

Graphical consoles offer a practical interface for medical acquisition assistance, allowing agents to work through the controls and visual feedback used by human operators. Reliable assistance requires linking on-screen anatomy to acquisition decisions that determine what image evidence becomes available next. We introduce MedImageOSWorld, a benchmark for evaluating this capability in simulated CT, MR, and ultrasound consoles. Using screenshots and mouse-and-keyboard actions, agents configure protocols, plan acquisitions, inspect the resulting images, and make corrective adjustments across seven capability levels, from console operation to feedback-driven control. Evaluation combines task-specific workflow checks with hidden anatomical ground truth to assess procedural completion and acquisition outcomes separately. A common evaluation protocol specifies episode conditions and interaction budgets, while recorded trajectories support analysis of how agents observe, act, and respond to acquisition feedback. Across eleven open-weight agents, success rates range from 3.0 to 25.0 on a 0-100 scale while workflow-progress rates reach 21.5-74.5: agents complete much of the console workflow but rarely acquire the intended anatomy. Success collapses between perception and planning, from 79-90% at the lowest three levels for the best agent to at most 12% for millimetre-level planning and 0% for closed-loop control among open-weight agents. Two proprietary agents reach success rates of about 34 and exceed the best open-weight agent mainly in acquisition quality (61 versus 42). MedImageOSWorld provides a controlled setting for studying whether general-purpose GUI agents can translate visual observations into effective medical acquisition decisions.

Comments11 pages, 3 figures. Supplementary material (8 pages) included as an ancillary file

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑