AI 中文总结
该研究针对临床视觉大语言模型的缺陷,提出BioMed-Agent-RL医学智能体,整合多模态元学习与强化学习等技术,在多基准实验中准确率达约73%,较现有模型提升约5%,为临床智能体系统建立新标准。
AI 中文摘要
当前临床视觉大语言模型(C-VLLMs)的进展大幅提升了数字诊断能力,但这些框架仍存在病灶噪声、模态错位、幻觉以及复杂临床病例中情境关联缺失的问题。此外,主流智能体系统通常依赖静态且不可适配的流程,缺乏复杂医学推理所需的通用性。为解决这些难题,我们提出BioMed-Agent-RL,一种统一的医学智能体,整合了自适应编排、策略以及基于奖励的强化学习(RL)模型以应用于生物医学场景。为确保可靠性,它调用了临床情境感知偏好优化(CPO)、直接偏好优化(DPO)以及带动态熵调控的分组相对策略优化(GRPO)。该流程采用多模态元学习方法,充当领域特定专家与人类判断合成器。该智能体通过迭代自适应RL方法,在各类临床模态(如X光)中自适应利用一组模型级专业能力,包括临床关联与推理器、病灶分割器、领域特定合成器。当误导性、冲突性视觉线索存在且专家建议有误时,该智能体学会合理综合信息并信任自身固有推理。我们在多个基准上开展了密集消融实验,结果显示该智能体显著优于现有最先进模型,如GPT-5,准确率最高达约73%,较当前基准提升约5%。因此,该框架为构建用于独立临床推理的事实性、可靠、鲁棒且类专家的智能体系统提出了新的标准。
英文摘要
The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often endure lesion noises, modality misalignment, hallucination, and missed contextual grounding in complex clinical cases. Moreover, prevailing agent systems usually depend on static and non-adaptable pipelines and lack the versatility necessary for complex medical reasoning. To resolve these difficulties, we present BioMed-Agent-RL, a unified medical agent that incorporates adaptive orchestration, policy, and reward-based reinforcement learning (RL) models for biomedical applications. To ensure reliability, it invokes clinical context-aware preference optimization (CPO), direct preference optimization (DPO), and group relative policy optimization (GRPO) with dynamic entropy regulation. This pipeline utilizes a multimodal meta-learning approach that operates as a field-specific expert and human judgment synthesizer. The agent adaptively utilizes a set of model-level expertise, such as clinical grounding and reasoner, lesion segmenter, and field-specific synthesizer, across various clinical modalities (e.g., X-ray) by utilizing an iterative and adaptive RL approach. The agent learns to seriously synthesize misleading, conflicting vision cues and trust in inherent reasoning, while specialist advice is faulty. An intensive ablation study is conducted across multiple benchmarks, and the agent significantly outperforms existing state of the art models, such as GPT-5, attaining up to ~73% accuracy (gain of ~5%) over contemporary baselines. As a result, the framework suggests a new standard for building factual, reliable, robust, and expert-like intelligent agent systems for independent clinical reasoning.