发表机构
Arionkoder LLC; Yatiris; UNICEN-CONICET(阿里昂科德有限责任公司; 亚蒂里斯; 国立中央大学-阿根廷国家科学技术研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种结合文献引导推理等的智能体式AI科学家工作流,可自动化医学影像模型开发流水线,在四个公开基准上取得竞争力结果,大幅降低相关工程工作量。
AI 中文摘要
为医学影像开发具有竞争力的深度学习基线仍是高度迭代的过程,需经过文献综述、实现、实验及专家优化。现有自动化方法通常仅优化孤立组件,如架构搜索或超参数调优,而非完整的基线开发流程。我们提出一种智能体式AI科学家工作流,结合文献引导推理、自动化代码生成及假设驱动实验,以生成医学影像挑战的竞争力基线模型。该框架在涵盖分割、分类及检测任务的四个公开基准上进行评估:所有任务中,实验流水线持续提升验证性能,取得竞争力排行榜结果,包括PUMA赛道(15支队伍)第6名、MILK10k(125支队伍)第31名;在MIDOG25上,所得模型还展现出跨扫描仪、肿瘤类型及物种的强领域泛化能力。在所有挑战中使用相同工作流,无需针对任务进行重新设计,我们证明基于技能、文献引导的智能体工作流可大幅减少开发竞争力医学影像基线所需的工程工作量。
英文摘要
Developing competitive deep learning baselines for medical imaging remains a highly iterative process requiring literature review, implementation, experimentation, and expert refinement. Existing automation approaches typically optimize isolated components, such as architecture search or hyperparameter tuning, rather than the complete baseline development process. We present an agentic AI Scientist workflow that combines literature-guided reasoning, automated code generation, and hypothesis-driven experimentation to generate competitive baseline models for medical imaging challenges. The framework is evaluated on four public benchmarks spanning segmentation, classification, and detection. Across all tasks, the Experimentation Pipeline consistently improves validation performance, achieving competitive leaderboard results, including 6th place on both PUMA tracks (15 teams) and 31st place on MILK10k (125 teams). On MIDOG25, the resulting model also demonstrates strong domain generalization across scanners, tumor types, and species. Using the same workflow across all challenges without task-specific redesign, we demonstrate that skill-based, literature-guided agentic workflows can substantially reduce the engineering effort required to develop competitive medical imaging baselines.
CommentsMICCAI 2026 Workshop AgenticMed