发表机构
University of California, Santa Cruz; NVIDIA(加州大学圣克鲁兹分校; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BaT系统,结合Stage Bank与BiCuRL方法,使医学研究智能体在AutoMedBench-Lite上得分大幅提升,BaT-9B性能优于Claude Opus 4.6。
AI 中文摘要
长程智能体正逐步实现代码、报告及研究制品等完整工作流的自动化。医学影像工作流多阶段且数据敏感,而专家轨迹稀缺且难以共享。结构化基准可通过阶段式评分标准定位失败,但标准后训练会在下一轮训练前丢弃这些诊断信息。本文提出Benchmark-as-Teacher(BaT),一种用于智能体后训练的递归自改进系统,包含两个关联组件:异步阶段库(Stage Bank)数据流水线与BiCuRL(双层课程强化学习)这一自改进后训练方法。Stage Bank在策略更新循环外合成内容隔离的训练状态;BiCuRL利用固定保留评估选择下一阶段课程,通过任务评分标准验证 rollout,使用GRPO更新策略,并将候选检查点返回评估。在AutoMedBench-Lite上,BaT-4B与BaT-9B的总体得分较其Qwen Instruct基线提升一倍以上;BaT-9B智能体总体得分达79.6,超过Claude Opus 4.6(搭配Claude Code)的77.5。
英文摘要
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.