arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BaT:构建具备阶段式评分标准的自演化医学研究智能体

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang

arXiv 2608.16211首次发表:更新:

发表机构

University of California, Santa Cruz; NVIDIA(加州大学圣克鲁兹分校; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BaT系统,结合Stage Bank与BiCuRL方法,使医学研究智能体在AutoMedBench-Lite上得分大幅提升,BaT-9B性能优于Claude Opus 4.6。

AI 中文摘要

长程智能体正逐步实现代码、报告及研究制品等完整工作流的自动化。医学影像工作流多阶段且数据敏感,而专家轨迹稀缺且难以共享。结构化基准可通过阶段式评分标准定位失败,但标准后训练会在下一轮训练前丢弃这些诊断信息。本文提出Benchmark-as-Teacher(BaT),一种用于智能体后训练的递归自改进系统,包含两个关联组件:异步阶段库(Stage Bank)数据流水线与BiCuRL(双层课程强化学习)这一自改进后训练方法。Stage Bank在策略更新循环外合成内容隔离的训练状态;BiCuRL利用固定保留评估选择下一阶段课程,通过任务评分标准验证 rollout,使用GRPO更新策略,并将候选检查点返回评估。在AutoMedBench-Lite上,BaT-4B与BaT-9B的总体得分较其Qwen Instruct基线提升一倍以上;BaT-9B智能体总体得分达79.6,超过Claude Opus 4.6(搭配Claude Code)的77.5。

英文摘要

Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑