深度认知价值函数用于乐观探索
Deep Epistemic Value Functions for Optimistic Exploration
浏览论文内容
中文总结 AI 辅助
针对深度认知价值函数在探索中的脆弱性,本文通过系统实证分析揭示其表示、传播和优化中的失败模式,提出DEVOTE算法,在无奖励探索和连续控制任务中更有效地到达新状态并获得更高回报。
中文摘要 AI 辅助
强化学习中的原则性探索要求智能体量化其认知不确定性并采取行动以消除之。价值函数上的不确定性为探索提供了自然的信号,然而现有的深度近似方法仍然脆弱且表现不一致。因此,核心挑战在于如何稳健地扩展这些思想。我们对深度认知价值函数中认知不确定性的表示、传播和优化进行了系统的实证研究,并在这三个轴向上发现了不同的失败模式。这些发现促使我们提出DEVOTE,一种无模型强化学习算法,它控制不确定性在观测数据之外的泛化方式,稳定其时间传播,并保持对由此产生的非平稳探索目标的适应性。在无奖励探索和具有挑战性的连续控制任务中,DEVOTE比强无模型和基于模型的探索基线更有效地到达新状态,并获得更高的任务回报。这些结果提供了证据,表明深度认知价值函数是实现可扩展、原则性探索的有前景的途径。
英文摘要
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
发表机构
- ETH Zürich(苏黎世联邦理工学院)
- Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。