发表机构
Hefei University of Technology; Nanjing University; Guangdong University of Finance and Economics; East China Normal University; Lenovo(合肥工业大学; 南京大学; 广东财经大学; 华东师范大学; 联想集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MedPrune提出动态剪除通信拓扑中节点和边的医学多模态多智能体协作框架,通过强化学习优化节点和边稀疏化,在提升推理能力与令牌效率的同时超越多智能体基线并保持强对抗鲁棒性。
AI 中文摘要
尽管医学多模态大语言模型(Med-MLLMs)推动了医学视觉问答(VQA)的发展,但现有受临床工作流启发的多智能体框架存在交互模式问题,并因冗余通信拓扑导致过高的计算开销。本文提出MedPrune,一种高效的医学多模态多智能体协作框架,通过动态剪除通信拓扑中的节点和边来增强推理能力和令牌效率。具体而言,我们首先将诊断过程形式化为异质通信图,其中节点代表来自不同科室的专家智能体,边捕获科室内部及跨科室的交互。在此基础上,我们引入两种稀疏化机制以实现自适应协作演化:(1)异质节点稀疏化,通过强化学习驱动的拓扑优化,消除与当前多模态问题无关的任务无关专家智能体;(2)异质边稀疏化,通过联合优化任务性能和拓扑复杂度,仅保留最具诊断意义的科室内部及跨科室连接。在全集和少样本训练设置下的大量医学VQA实验证明,MedPrune优于多智能体基线,并在提升令牌效率的同时具有强大的对抗鲁棒性。
英文摘要
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.