arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32984cs.CV

ReVision3D:用于三维医学感知的归因引导递归自我改进

ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception

Ho Hin Lee, Yuyin Zhou, Yannan Yu, Shi Gu, Yifan Wu

首次发表
浏览论文内容

中文总结 AI 辅助

ReVision3D通过归因引导的递归自我改进,利用带标注三维体积定位视觉证据丢失点,迭代优化医学影像智能体,在腹部CT上显著提升小病灶检测性能。

中文摘要 AI 辅助

递归自我改进(RSI)为克服当前医学影像智能体视觉能力有限的问题提供了一条有前景的路径。然而,将RSI应用于体积成像仍然困难:失败可能源于采集、感知、训练方案或下游推理,而自我生成的反馈和记录的轨迹几乎无法指导应更改哪个组件。我们提出了ReVision3D,一个RSI系统,利用带有空间标注的三维体积来确定视觉证据在何处丢失,并递归改进相应的视觉能力。一个冻结的语言模型设计者提出对采集、感知、训练或推理的修订,而验证器和系统级目标保持不变。我们的关键见解是,带标注的体积构成了视图渲染和空间验证的精确重放世界:可以按需渲染未访问的视图,并且可以直接对照参考掩膜检查局部预测。这种基于地面的反馈指导有针对性的修订,而只有那些改进超过测量到的种子噪声的更改才会被保留。每次接受的更改都会触发新的归因,使得主导瓶颈可以在多轮中转移。在腹部CT上,归因识别出感知是主导的剩余限制。修订该级别使ReVision3D在每患者少于0.4个假阳性的情况下实现79%的肝脏召回率和83%的肾脏召回率,优于所评估的冻结多模态基础模型,在小病灶上增益最大。

英文摘要

Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-generated feedback and logged trajectories provide little guidance on which component should change. We introduce ReVision3D, an RSI system that leverages 3D volumes with spatially grounded annotations to determine where visual evidence is lost and recursively improve the corresponding visual capability. A frozen language-model designer proposes revisions to acquisition, perception, training, or inference, while the verifier and system-level objective remain fixed. Our key insight is that an annotated volume forms an exact replay world for view rendering and spatial verification: unvisited views can be rendered on demand, and localized predictions can be checked directly against reference masks. This grounded feedback directs targeted revision, while only changes that improve beyond measured seed noise are retained. Each accepted change triggers renewed attribution, allowing the dominant bottleneck to shift across rounds. On abdominal CT, attribution identifies perception as the dominant remaining limitation. Revising that level enables ReVision3D to achieve 79% liver recall and 83% kidney recall at under 0.4 false positives per patient, outperforming the evaluated frozen multimodal foundation models, with the largest gains on small lesions. Our framework is public available with interactive demo in this project page: \url{https://leeh43.github.io/ReVision3D/}

补充信息

↑