arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从人类视角理解:用于交互式自我中心医学图像分割的多智能体系统

Understanding From Human Perspective: A Multi-agent System for Interactive Egocentric Medical Image Segmentation

Rongjun Ge, Dongyang Wang, Heng Zhu, Zhirui Li, Yang Chen, Yuting He

arXiv 2607.17341首次发表:更新:

发表机构

School of Instrument Science and Engineering, Southeast University; School of Computer Science and Engineering, Southeast University; Sichuan University-Pittsburgh Institute, Sichuan University; Department of Biomedical Engineering, Case Western Reserve University(东南大学仪器科学与工程学院; 东南大学计算机科学与工程学院; 四川大学匹兹堡学院; 凯斯西储大学生物医学工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究交互式自我中心医学图像分割面临的挑战,提出EgoMed-Agent多智能体系统,通过目标确认和定位引导传播工作流程从人类视角理解目标,实验显示该系统平均Dice远超基线,效果显著。

AI 中文摘要

交互式自我中心医学图像分割(IEMIS)在智能眼镜辅助医学图像审查中发挥重要作用,可从临床医生的自我中心视角分割医学目标,为审查提供视觉证据并支持细粒度分析和临床决策。但来自用户自我中心视角的指令和视频带来挑战,包括语义模糊和视觉变异性。本文提出EgoMed-Agent多智能体系统,通过目标确认工作流程和定位引导传播工作流程从人类视角理解目标。实验表明,EgoMed-Agent平均Dice达到71.34%,远超最佳文本提示基线(11.70%)。

英文摘要

Interactive egocentric medical image segmentation (IEMIS) plays an important role in smart-glasses-assisted medical image review, segmenting the medical targets a clinician refers to from their egocentric view. Once it succeeds, the object-level visual evidence it provides strengthens the review and underpins fine-grained analysis and clinical decision-making. However, the instruction and the video both come from the user's egocentric perspective, which poses two challenges. (1) Semantic ambiguity leaves the model unable to confirm the user-intended target. (2) Visual variability makes the segmentation jump from frame to frame. In this paper, we propose EgoMed-Agent, a multi-agent system that understands the target from the human perspective through two workflows. (1) The \textit{Target Confirmation Workflow} grounds the instruction against candidate targets with a reliability score, confirming the target when the grounding is reliable and asking the user to clarify when it is not, thereby confirming the segmentation target. (2) The \textit{Localization-Guided Propagation Workflow} couples mask propagation with per-frame target localization, using the localized target to correct the propagated mask whenever the two diverge, so the segmentation stays on the target across the egocentric video. Extensive experiments show that EgoMed-Agent reaches 71.34\% average Dice, far above the best text-prompted baseline (11.70\%). Our code is available at \href{https://github.com/wdyyyyyy/EgoMed-Agent}{our project page}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑