arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DoctorAgents:用于针对小型临床时序数据迭代优化自动机器学习流水线的智能体框架

DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data

Ruilin Wang, Bo-Hong Wang, Elizabeth Kourbatski, Jun Bai, Hegang Chen, Ziyang Song, Gilles Boire, Marie Hudson, Yue Li

arXiv 2608.05375首次发表:更新:

发表机构

McGill University; Mila – Quebec AI Institute; CIUSSSE-CHUS; University of Sherbrooke(麦吉尔大学; 米拉-魁北克人工智能研究所; 谢布鲁克大学医学中心综合大学健康和社会服务中心; 谢布鲁克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DoctorAgents是一种用于小型临床时序数据的智能体框架,通过LLM智能体以推理驱动方式迭代优化AutoML流水线,在多种临床任务上性能优于现有AutoML基线且可解释性更强。

AI 中文摘要

临床机器学习(ML)有潜力支持高风险医疗决策,但可靠部署常受限于数据稀缺、异质性及时序复杂性。为这类数据开发有效ML流水线耗时且易出错,现有自动机器学习(AutoML)系统仅部分解决该挑战,因其大多依赖预定义空间的暴力搜索,缺乏明确推理与记忆。为此,我们将小型临床数据的AutoML从穷举搜索重构为推理驱动的优化,提出DoctorAgents——一种智能体AI框架,通过专门的大型语言模型(LLM)智能体(含生成、验证、优化功能)自主构建并优化端到端ML流水线。DoctorAgents通过文本梯度下降反向传播自然语言反馈,在无需穷举搜索的情况下执行针对性更新。在多种临床任务上的实验表明,DoctorAgents始终优于现有AutoML基线,同时生成更具可解释性的任务特定表示。

英文摘要

Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this challenge because they largely rely on brute-force search over predefined spaces and lack explicit reasoning and memory. We therefore reformulate AutoML for small clinical data from exhaustive search to reasoning-driven refinement. We propose DoctorAgents, an agentic AI framework that autonomously constructs and optimizes end-to-end ML pipelines through specialized large language model (LLM) agents for generation, validation, and refinement. DoctorAgents backpropagates natural-language feedback through textual gradient descent to perform targeted updates without exhaustive search. Experiments across diverse clinical tasks show that DoctorAgents consistently outperforms established AutoML baselines while producing more interpretable task-specific representations.

Comments34 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑