MedVA:一种用于医学体数据可视化的端到端神经符号智能体系统
MedVA: An End-to-End Neuro-Symbolic Agentic System for Medical Volume Visualization
查看机构详情
- Gachon University(嘉泉大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
MedVA是用于医学体数据可视化的端到端神经符号智能体系统,通过三个互补智能体解决MLLM推理在ROI识别和可视化优化中的不足,实验和用户研究验证了其有效性。
中文摘要 AI 辅助
医学体数据可视化需要根据给定的临床意图选择感兴趣区域(ROIs)并仔细控制其相对视觉强调。在传统工作流程中实施这些决策需要大量的临床和可视化专业知识,且通常涉及反复试错的优化过程。最近的智能体系统引入了自然语言交互和自主可视化操作,但在整个工作流程中主要依赖基于多模态大语言模型(MLLM)的推理。尽管MLLM编码了广泛的医学知识并提供强大的推理能力,但这种推理对于医学体数据可视化可能并非最优,可能导致对用户请求的临床解释不完整以及ROI识别和可视化优化不可靠。在本工作中,我们提出了MedVA,一种用于医学体数据可视化的端到端神经符号智能体系统,通过三个互补的智能体解决上述局限性。神经符号意图制定智能体通过对既有临床知识进行符号推理来细化基于MLLM的自然语言请求解释,从而提供比仅使用MLLM推理更完整、更具临床依据的ROI规格。多模型ROI识别智能体通过利用互补的大规模预训练医学分割模型,直接在原始体数据中识别语义指定的ROI。目标驱动的可视化优化智能体使用基于体数据的可见性目标显式评估原始体数据中的ROI可见性和遮挡情况。跨多种医学数据集和交互场景的广泛智能体级和系统级评估支持了各个智能体的有效性。一项形成性用户研究进一步表明,该系统在不同专业水平的用户中具有较高的可用性和实用价值。
英文摘要
Medical volume visualization requires selecting regions of interest (ROIs) and carefully controlling their relative visual emphasis according to a given clinical intent. Implementing these decisions in conventional workflows demands substantial clinical and visualization expertise and often involves trial-and-error optimization. Recent agentic systems have introduced natural-language interaction and autonomous visualization operations but largely rely on MLLM-based inference throughout the workflow. Although MLLMs encode broad medical knowledge and provide strong reasoning capabilities, such inference may be suboptimal for medical volume visualization, potentially leading to clinically incomplete interpretations of user requests and unreliable ROI identification and visualization optimization. In this work, we present MedVA, an end-to-end neuro-symbolic agentic system for medical volume visualization that addresses these limitations through three complementary agents. The neuro-symbolic intent formulation agent refines MLLM-based interpretations of natural-language requests through symbolic reasoning over established clinical knowledge, which provides more complete, clinically grounded ROI specifications than MLLM-only reasoning. The multi-model ROI identification agent directly identifies semantically specified ROIs in the original volume by leveraging complementary large-scale pretrained medical segmentation models. The objective-driven visualization optimization agent explicitly evaluates ROI visibility and occlusion in the original volume using a volume-based visibility objective. Extensive agent-level and system-level evaluations across diverse medical datasets and interaction scenarios support the effectiveness of the individual agents. A formative user study further indicates high usability and practical value among users with different levels of expertise.