arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedZERO:通过受控知识积累实现开放式医学推理的自进化智能体

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Xilin Dang, Weilin Ruan, Xue Yang, Jinghao Wang, Xiaowei Hu, Jinpeng Li, Pheng-Ann Heng

arXiv 2610.08327首次发表:更新:

发表机构

The Chinese University of Hong Kong; Zhongshan Hospital, Fudan University; South China University of Technology; Chinese Academy of Sciences(香港中文大学; 复旦大学附属中山医院; 华南理工大学; 中国科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MedZERO提出自进化框架,结合检查器与推理器,通过受控知识积累实现开放式医学推理,在多个基准上显著超越基线。

AI 中文摘要

大型语言模型(LLMs)在医学问答和临床推理方面展现出潜力,但其改进仍受限于静态参数知识和昂贵的专家监督。自进化智能体提供了一种有前景的替代方案,使模型能够通过迭代任务生成和问题解决来改进。然而,大多数现有的自进化方法是为数学和编程等易于验证的领域设计的,在这些领域中,解决方案可以通过精确答案或可执行程序来检查。医学推理从根本上不同:它是开放式的、知识密集型的,且通常只能部分验证。我们提出了MedZERO,一个用于开放式医学推理的自进化框架。MedZERO将生成前沿医学问题-选项对的检查器(Examiner)与通过基于证据的多轮推理和外部知识工具解决这些问题的推理器(Reasoner)相结合。为了支持可靠、持续的改进,MedZERO采用受控知识积累,在推理中维护临时探索性知识和精选的持久性知识。我们在五个公共医学推理基准上,使用4B和8B规模的基础模型,在开放式评估下评估了MedZERO。在所有设置中,MedZERO始终优于底层基础模型和先前的自进化基线,相较于次优的自进化基线,平均准确率提升最高达13.7个百分点。

英文摘要

Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.

Commentsaccepted by NIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑