CT-Flow: 通过模型上下文协议服务器协调CT解释工作流
CT-Flow: Orchestrating CT Interpretation Workflow with Model Context Protocol Servers
- Qing Yuan Research Institute, Shanghai Jiao Tong University, Shanghai, China(清元研究院,上海交通大学,上海,中国)
- Shanghai Innovation Institute, Shanghai, China(上海创新研究院,上海,中国)
- Sensetime Research, Shanghai, China(senseTime研究院,上海,中国)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CT-Flow通过模型上下文协议实现3D CT解释工作流的动态协调,提升诊断准确性和工具调用效率。
AI中文摘要:
近年来,大型视觉-语言模型(LVLMs)在多模态放射学推理方面的进展显示出强大的潜力,特别是在诊断视觉问答(VQA)和放射学报告生成等任务中。然而,大多数现有的3D CT分析方法主要依赖于静态的单次推理。在实践中,临床解释是一个动态的、工具介导的工作流程,放射科医生会反复查看切片,并使用测量、放射组学和分割工具来细化发现。为了弥合这一差距,我们提出了CT-Flow,一种用于互操作性体积解释的代理框架。通过利用模型上下文协议(MCP),CT-Flow从封闭式推理转向开放的、工具感知的范式。我们整理了CT-FlowBench,这是第一个大规模指令微调基准,专门针对3D CT工具使用和多步推理。在此基础上,CT-Flow作为临床协调器,能够将复杂的自然语言查询分解为自动工具使用序列。在CT-FlowBench和标准3D VQA数据集上的实验评估表明,CT-Flow在诊断准确性上超越了基线模型,提高了41%,并在自主工具调用中达到了95%的成功率。这项工作为将自主、代理智能整合到现实世界的临床放射学中提供了可扩展的基础。
英文摘要:
Recent advances in Large Vision-Language Models (LVLMs) have shown strong potential for multi-modal radiological reasoning, particularly in tasks like diagnostic visual question answering (VQA) and radiology report generation. However, most existing approaches for 3D CT analysis largely rely on static, single-pass inference. In practice, clinical interpretation is a dynamic, tool-mediated workflow where radiologists iteratively review slices and use measurement, radiomics, and segmentation tools to refine findings. To bridge this gap, we propose CT-Flow, an agentic framework designed for interoperable volumetric interpretation. By leveraging the Model Context Protocol (MCP), CT-Flow shifts from closed-box inference to an open, tool-aware paradigm. We curate CT-FlowBench, the first large-scale instruction-tuning benchmark tailored for 3D CT tool-use and multi-step reasoning. Built upon this, CT-Flow functions as a clinical orchestrator capable of decomposing complex natural language queries into automated tool-use sequences. Experimental evaluations on CT-FlowBench and standard 3D VQA datasets demonstrate that CT-Flow achieves state-of-the-art performance, surpassing baseline models by 41% in diagnostic accuracy and achieving a 95% success rate in autonomous tool invocation. This work provides a scalable foundation for integrating autonomous, agentic intelligence into real-world clinical radiology.