Branch2Skill:基于推理树的高效技能演化方法
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
浏览论文内容
中文总结 AI 辅助
Branch2Skill是将推理树转化为密集监督信号的框架,通过蒙特卡洛树搜索提取推理证据并提炼更新,在六个基准测试中提升了任务性能与技能演化效率,以GPT 5.5为目标模型时token用量比SkillOpt少73.2%
中文摘要 AI 辅助
技能演化通过随时间推移的反馈提升智能体的技能,失败的轨迹往往能揭示不完整或误导性行为,从而提供有价值的信号。然而,现有方法主要依赖单一轨迹,早期的推理错误会通过后续步骤传播,削弱用于技能优化的反馈。因此,提升技能需要多轮的rollout(轨迹生成)、诊断和更新循环,产生大量token开销。为解决这一挑战,我们提出Branch2Skill,这是一个将单一推理树转化为密集监督信号的高效技能演化框架。对于每个任务或问题,Branch2Skill在固定预算下执行蒙特卡洛树搜索以获得多样化的推理轨迹,随后将精英路径与具有相同前缀的兄弟备选路径进行比较,提取关于应保留、修正或避免哪些推理模式的逐步证据。最后,Branch2Skill将多步证据提炼为可复用的更新,使得一个推理树能为多个推理步骤提供监督,减少对重复rollout-更新循环的需求。在涵盖推理和智能体任务的六个基准测试中,Branch2Skill在提升任务性能的同时,还提高了技能演化的效率。例如,以GPT 5.5作为目标模型,Branch2Skill的token使用量比SkillOpt少73.2%,同时实现了更优的性能。这些结果表明,推理树不仅能支持更有效的轨迹搜索,还能提供更丰富的监督信号,实现更高效的技能提升。代码将被公开。
英文摘要
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. However, existing methods mainly rely on single trajectories, where early reasoning errors can propagate through subsequent steps and weaken the feedback available for skill refinement. Consequently, improving skills requires repeated cycles of rollout, diagnosis, and update, incurring substantial token costs. To address this challenge, we introduce Branch2Skill, an efficient framework that transforms a single reasoning tree into dense supervision for skill evolution. For each task or problem, Branch2Skill performs Monte Carlo tree search under a fixed budget to obtain diverse reasoning trajectories, then compares an elite path with sibling alternatives sharing the same prefixes to extract step-wise evidence about which reasoning patterns to retain, revise, or avoid. Finally, Branch2Skill distills multi-step evidence into reusable updates, allowing one reasoning tree to provide supervision across multiple reasoning steps and reducing the need for repeated rollout-update cycles. Across six benchmarks covering reasoning and agentic tasks, Branch2Skill consistently improves task performance while enhancing skill evolution efficiency. For example, with GPT 5.5 as the target model, Branch2Skill uses 73.2% fewer tokens than SkillOpt, while achieving superior performance. These results demonstrate that reasoning trees can support not only more effective trajectory search, but also richer supervision for more efficient skill improvement. Code will be published.