CHART:一种用于搜索智能体的挽具轮换课程训练方法
CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents
浏览论文内容
中文总结 AI 辅助
针对搜索智能体在挽具更新后行为失效的问题,提出CHART课程式挽具轮换训练方法,通过动态轮换挽具保持奖励差距,使模型在所有挽具上学会并行搜索并提升迁移性能。
中文摘要 AI 辅助
搜索智能体通常在单一挽具(即系统提示等运行环境配置)下进行训练。然而,一旦智能体部署到实际应用中,其挽具经常会被更新(例如重写系统提示)以适应生产需求。这暴露了后训练智能体的一个脆弱性:由于学习到的行为与其训练挽具纠缠在一起,即使挽具更新不改变任务本身,也可能无法触发该行为。我们训练一个搜索智能体执行并行搜索,这是一种同时提升搜索效率和性能的常用策略。我们发现,在固定挽具下训练会使行为具有挽具局部性,即过度拟合该挽具的表面形式:当挽具改变时,模型会退回到串行搜索。一个直观的修复方法是挽具增强,但仅仅在更多挽具上训练并不能解决该问题。GRPO从同一问题的并行和串行rollout之间的奖励差距中学习:小的挽具池会过早地使该差距饱和,而大的挽具池则会稀释每个挽具的信号,使得任何挽具都无法巩固学习。因此,我们提出了课程式挽具轮换训练(CHART),这是一种轮换课程,让搜索智能体逐步在多个挽具间巩固并行搜索能力。在每次周期性评估中,CHART会“毕业”那些预期行为已被学会的挽具,并用仍可学习的挽具替换它们,从而在整个训练过程中保持奖励差距的活跃。从相同的挽具池出发,CHART使模型在所有挽具上都学会了并行搜索,而静态增强最多只能在一半的挽具上成功。该行为还能迁移到未见过的挽具上:CHART在89%的未见过的轮次上实现了并行化,而静态池最多只有5%。它还能进一步迁移到新的问答任务和搜索环境,在pass@1上比最佳静态池提高了5.6个百分点。最后,经过CHART训练的智能体从元挽具搜索中获得的收益比基线更大。
英文摘要
Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned behavior is entangled with its training harness, even a harness update that leaves the task unchanged can fail to elicit the behavior. We train a search agent to perform parallel search, a popular strategy for improving both search efficiency and performance. We find that training under a fixed harness makes the behavior harness-local, overfit to that harness's surface form: when the harness changes, the model falls back to serial search. An intuitive fix is harness augmentation, but simply training on more harnesses does not resolve the problem. GRPO learns from the reward gap between parallel and serial rollouts of the same question: a small harness pool saturates that gap early, while a large pool dilutes the per-harness signal too thinly for any harness to consolidate. We therefore propose Curriculum HArness Rotation Training (CHART), a rotating curriculum that lets a search agent gradually consolidate parallel search across harnesses. At each periodic evaluation, CHART "graduates" the harnesses whose expected behavior is learned and replaces them with still-learnable ones, keeping the reward gap alive throughout training. Starting from the same harness pool, CHART makes the model learn parallel search on all harnesses, whereas static augmentation succeeds on at most half of them. The behavior also carries to held-out harnesses: CHART parallelizes on 89% of held-out turns, against at most 5% for the static pools. It further transfers to a new QA task and search environment, improving pass@1 by 5.6pp over the best static pool. Finally, CHART-trained agents benefit more from meta-harness search than baselines.
发表机构
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。