发表机构
Institute of Automation; Chinese Academy of Sciences(自动化研究所; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过理论和实验解释了选项-评论中增加选项提升性能的原因:终止规则无贡献,策略坏死导致状态锁定,额外选项通过降低联合失败概率(从59%降至4%)来提升性能。
AI 中文摘要
选项-评论算法学习选项:子策略以及一个学习规则,用于决定每个子策略何时将控制权交回。其最显著的结果是,随着选项数量的增加,性能会提升。我们通过理论和实验解释了这一结果。首先,选项-评论通过最大化回报来学习的终止规则并未做出贡献。当终止测试和选择选项的策略读取相同的值时,测试在每一步都会触发,因此学习到的规则与始终终止相同。当该策略进行探索而测试不进行探索时(如选项-评论本身的情况),该规则可能阻碍探索;存在一些实例,其中它遭受$\Omega(T)$的遗憾,而始终终止则保持在$O(\log T)$。强制每一步终止会保持选项数量曲线不变。其次,选项内部的策略几乎不进行探索,因此一个状态会锁定到第一个看起来不错的动作,并且永远不会再更新。我们将此命名为策略坏死,给出了一个状态级别的测试,并发现典型选项中有五分之三的状态是坏死的。恢复探索可以修复这些状态,然后一个选项就能解决任务。第三,额外的选项不会改进任何选项;下降的是所有选项在同一状态失败的概率,从$59\\%$降至$4\\%$,性能跟随这一联合量。
英文摘要
Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. Its headline result is that performance improves as options are added. We explain that result, with theory and experiment. First, the termination rule option-critic learns by maximising return contributes nothing. When the termination test and the policy that picks options read the same values, the test fires at every step, so the learned rule is identical to always terminating. When that policy explores and the test does not, as in option-critic itself, the rule can block the exploration; there are instances where it suffers $Ω(T)$ regret while always terminating holds to $O(\log T)$. Forcing termination at every step leaves the option-count curve intact. Second, the policy inside an option barely explores at all, so a state locks onto the first action that looked good and never updates again. We name this policy necrosis, give a state-level test for it, and find three fifths of states necrotic in a typical option. Restoring exploration repairs those states, and one option then solves the task. Third, extra options improve no option; what falls is the chance that all of them fail in the same state, from $59\\%$ to $4\\%$, and performance follows that joint quantity.
CommentsFirst two authors contributed equally