arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于验证的上下文工程实现全自动医学影像代码生成

Towards Fully Automated Medical Imaging Code Generation via Validation-based Context Engineering

Zixiao Zhao, Jing Sun, Zhe Hou, Cheng-Hao Cai, Qian Liu, Mengze Li, Zijian Zhang, Jin Song Dong

arXiv 2608.29016首次发表:更新:

发表机构

University of Auckland; Griffith University; Suzhou Industrial Park Monash Research Institute of Science and Technology; Beijing Institute of Technology; National University of Singapore(奥克兰大学; 格里菲斯大学; 苏州工业园区莫纳什科技研究院; 北京理工大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AutoMedImg多智能体框架,通过规划与编码阶段的多阶段验证及自动上下文工程,实现零人工干预的全自动医学影像代码生成,在六个数据集上分割任务Dice最高0.90、分类准确率99%。

AI 中文摘要

大语言模型(LLMs)在小规模常规应用开发的代码生成中展现出巨大潜力,但应用于医学图像处理等复杂领域特定任务时仍存在局限:通用模型缺乏明确的领域知识和可靠的验证机制以确保正确性,通常需要大量人工干预才能生成可靠的处理流程。为解决这些局限,本文提出AutoMedImg,一种用于全自动医学图像处理代码生成的多智能体框架。AutoMedImg分两个阶段协调专用智能体:规划阶段执行数据集分析与架构设计,同时进行语义和形式验证;编码阶段并行生成模块,并完成静态检查、执行测试与组装验证。这种多阶段验证可减轻生成过程中的错误传播,结合领域知识库、共享内存和验证反馈的全面自动上下文工程可自动构建上下文而无需人工提示。跨项目自适应流程合成机制还能积累已验证的流程,并根据项目相似性为新任务检索成熟组件,通过跨项目学习提升生成效率。在六个不同且公认的医学影像数据集上,对五种骨干LLMs进行的广泛评估表明,AutoMedImg实现了零人工干预,分割任务的Dice分数最高达0.90,分类任务准确率达99%。

英文摘要

Large language models (LLMs) have demonstrated considerable promise in program generation for small-scale and conventional application development; however, they remain limited when applied to complex, domain-specific tasks such as medical image processing. General-purpose models lack explicit domain knowledge and robust validation mechanisms to ensure correctness, often requiring substantial human intervention to produce reliable processing pipelines. To address these limitations, we propose AutoMedImg, a multi-agent framework for fully automated medical image processing code generation. AutoMedImg orchestrates specialised agents across two phases: a Planning Phase that performs dataset analysis and architecture design with semantic and formal verification, and a Coding Phase that generates modules in parallel with static checking, execution testing, and assembly validation. This multi-stage validation mitigates error propagation throughout generation, while comprehensive auto-context engineering combining domain-specific knowledge bases, shared memory, and validation feedback automates context construction without manual prompting. A cross-project adaptive pipeline synthesis mechanism further accumulates validated pipelines and retrieves proven components for new tasks based on project similarity, enhancing generation efficiency through cross-project learning. Extensive evaluation across six diverse and well-established medical imaging datasets with five backbone LLMs demonstrates that AutoMedImg achieves zero human intervention, with Dice scores of up to 0.90 for segmentation tasks and 99% accuracy for classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑