arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15700cs.AI

搜索策略与学习策略的自适应混合

Adaptive Mixing of Policies from Searching and Policies from Learning

Gavin B. Rens

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对强化学习中搜索耗时问题,提出Flexer架构自适应混合神经网络与MCTS策略,在三类玩具符号问题实验中性能优于AlphaZero、DQN和ADP。

中文摘要 AI 辅助

背景:通过搜索或规划生成的训练目标的蒸馏已被证明对强化学习有用,但搜索可能耗时极长。目标:不再每次都执行相同深度的搜索(通常是固定步数周期),而是根据策略网络先验的质量成比例地降低搜索深度。方法:我们描述了Flexer,一种架构,每一步都混合神经网络策略和蒙特卡洛树搜索(MCTS)策略。当网络的策略模仿误差和环境模型的方差增大时,混合因子会偏向MCTS策略。结果:在三个玩具符号问题的一些实验中,Flexer的性能优于AlphaZero的一个版本以及DQN和ADP。

英文摘要

Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can take exceedingly long. Objectives: Rather than perform search to the same depth every time (typically at a fixed period of steps), reduce the search depth proportionally to the quality of the policy network priors. Methods: We describe Flexer, an architecture that, for each step, mixes the policy from a neural network and the policy from Monte Carlo tree search. The mixing factor favors the MCTS policy as the policy imitation error of the network and the environment models' variance increases. Results: Flexer outperforms a version of AlphaZero (and DQN and ADP) for some experiments on three toy symbolic problems.

发表机构

  • Stellenbosch University(斯泰伦博斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑