ActiveMedAgent:面向多模态医疗诊断的成本感知轨迹学习
ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis
浏览论文内容
中文总结 AI 辅助
ActiveMedAgent是将成本感知序贯逻辑引入多模态医疗AI的框架,通过轨迹学习训练MLP控制器,在三个基准测试中优于无引导采集和全模态基准,还发现信息过载效应,少通道采集也能正确诊断。
中文摘要 AI 辅助
临床诊断本质上是序贯过程:临床医生仅在预期额外证据可解决诊断不确定性时,才会从低成本检查升级到高成本检查。我们提出ActiveMedAgent,这一框架将这种成本感知的序贯逻辑引入多模态医疗AI。给定一个已冻结、可通过API访问的视觉语言模型,ActiveMedAgent会跟踪候选诊断的概率分布,并对每次信息采集按其每一步诊断效用减去成本进行评分。随后,一个轻量级MLP控制器会基于这些评分后的轨迹进行离线训练,学习何时请求额外证据、何时做出诊断决策。在三个常用基准测试中,基于轨迹的策略学习始终优于无引导采集和全模态基准。值得注意的是,我们发现了信息过载效应:在175个案例中,该智能体用更少的信息通道即可做出正确诊断,而全模态基准却失败了,这表明学习该省略什么与学习该采集什么同等重要。
英文摘要
Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
发表机构
- Washington University in St. Louis(圣路易斯华盛顿大学)
- Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)
- Arizona State University(亚利桑那州立大学)
- Yale University(耶鲁大学)
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。