发表机构
The Chinese University of Hong Kong; Institute of Automation, Chinese Academy of Sciences; Microsoft Research; Lehigh University(香港中文大学; 中国科学院自动化研究所; 微软研究院; 里海大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对医学成像模型开发困难的问题,提出AMID自主多代理框架。该框架含数据条件方法规划和验证引导两阶段优化,经20个医学成像任务验证,其性能优于通用MLE系统,能将特定任务模型开发转变为高效可审计的代理工作流程。
AI 中文摘要
大语言模型(LLM)代理开始通过结合规划、代码执行、调试和经验反馈来自动化机器学习工程(MLE)。将此能力应用于医学成像仍很困难,因为每个任务都有特定模态的实验以及对验证协议和预测工件的严格要求。本文介绍了AMID,一个用于医学成像模型开发的自主多代理框架。AMID首先提出数据条件方法规划,将粗略的任务级搜索空间细化为基于特定任务数据分析和可运行医学成像资源的可执行、可并行化方法路径。然后开发验证引导的两阶段优化,从对不同方法路径的广泛早期探索转向对有前景候选者的选择性利用,同时在整个优化过程中严格验证验证协议、指标计算和预测工件。在跨越不同模态和预测类型的20个医学成像挑战任务中,AMID优于评估的通用MLE系统,在几个任务上接近或匹配强大的人工设计挑战解决方案。这些结果表明,AMID可以将特定任务的医学成像模型开发从定制的手动工程转变为跨异构任务生成高性能和可审计模型工件的代理工作流程。
英文摘要
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.
Comments18 Pages