MedRouter:揭示医学大语言模型间的知识差异以实现基于路由的推理
MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
中文总结 AI 辅助
MedRouter通过嵌入多标签路由器和两阶段SCALE训练框架,组合专科医学大模型,在八个基准上平均准确率提升8%,实现更全面的医学推理。
中文摘要 AI 辅助
医学问答涵盖多种专科和模态,单个医学大语言模型(LLM)在不同任务和领域展现出各自独特的优势。这种异质性表明,组合多个专科模型可能比依赖任何单一模型能够更广泛地覆盖医学问题。然而,现有的LLM路由方法主要旨在平衡答案质量和推理成本,尚未探索如何利用专科能力的差异来改进医学推理。在本文中,我们引入了MedRouter,一个智能体系统,它使用基于嵌入的多标签路由器来选择并查询专科LLM,然后将它们的响应传递给生成器以产生最终答案。我们进一步提出了SCALE(专科能力感知学习),一个两阶段的训练框架,首先使用专科正确性监督训练路由器,然后通过强化学习优化其选择。第二阶段使用性能增益奖励(PGR),该奖励衡量专科信息相对于不包含该信息时对生成器答案正确性的影响。在八个基于文本和多模态的医学问答基准上的实验表明,MedRouter在平均准确率上比最强的路由基线高出8%。我们对专科输出的分析进一步揭示了不同的优势和互补的问题级覆盖,这激励了学习路由以组合这些能力,实现更全面的医学推理。
英文摘要
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any single model. However, existing LLM routing methods primarily seek to balance answer quality and inference cost, leaving open how to exploit differences in specialist competence to improve medical reasoning. In this paper, we introduce MedRouter, an agentic system that uses an embedding-based multi-label router to select and query specialist LLMs, then passes their responses to a generator to produce the final answer. We further propose SCALE (Specialist Competence-Aware Learning), a two-stage training framework that first trains the Router with specialist correctness supervision and then optimizes its selections through reinforcement learning. The second stage uses a Performance Gain Reward (PGR) that measures how specialist information affects the generator's answer correctness relative to answering without that information. Experiments on eight text-based and multimodal medical QA benchmarks show that MedRouter outperforms the strongest routing baseline by 8% in average accuracy. Our analysis of specialist outputs further reveals distinct strengths and complementary question-level coverage, motivating learned routing to combine these capabilities for more comprehensive medical reasoning.