arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05111cs.LG

奖励结构塑造强化学习中情景探索与神经记忆的交互作用

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

Jai Malegaonkar, Rohan Patil, Henrik I. Christensen

AI总结:

该研究通过对照实验揭示,强化学习中情景探索奖励与神经记忆的交互受奖励结构调控,二者互补,奖励诱导暴露、记忆将其转化为回报,还形式化了奖励稀疏性相关概念。

AI中文摘要:

在部分可观测强化学习中,智能体面临双重瓶颈:它们必须进行探索以遇到有奖励的状态,并将该经验保留在记忆中以优化其策略。传统上,探索奖励和记忆架构是单独评估的,忽略了它们之间的交互作用,且标准的稀疏奖励概念将时间信号密度与奖励实际监督的内容混为一谈。我们开展了一项对照研究,在三种记忆内容获取方式不同的环境中,交叉对比情景探索奖励与多种神经记忆架构。相同的奖励信号产生了三种不同的交互模式:在必须主动发现并无监督保留记忆内容的场景中,它放大了架构容量差异;在记忆内容一旦被找到就是单一奖励监督线索的场景中,它将架构均衡至共享上限;在观测流完全按计划安排的场景中,它无作用。对照奖励操作验证这些模式跟踪的是奖励结构而非密度:仅当密集奖励直接监督所需的潜在记忆时,它才会抵消奖励;对探索性动作施加小的可避免惩罚(不改变最优解)会导致策略收敛到次优的稳态,而任一奖励都能解决该问题。我们随后用锚定观测的奖励机器形式化奖励稀疏性,将结构稀疏性(自动机无需任务所需历史即可复制回报)与潜在稀疏性(单步奖励对局部探索性动作定价不当)区分开来;所得术语按每个任务暴露的保留负担对三种机制进行组织。这些结果共同表明,探索与记忆是互补而非替代关系:奖励诱导暴露,且仅记忆能将暴露转化为回报。

英文摘要:

In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content of memory is acquired. An identical bonus signal yields three distinct interaction patterns: it amplifies architectural capacity differences where memory content must be actively discovered and retained unsupervised; equalizes architectures to a shared ceiling where the content, once sought out, is a single reward-supervised cue; and is null where the observation stream is purely scheduled. Controlled reward manipulations verify that these patterns track reward structure rather than density: a dense reward neutralizes a bonus only if it directly supervises the required latent memory, and a small avoidable penalty on exploratory actions (leaving the optimum unchanged) induces policy convergence to suboptimal stationary states, which either bonus resolves. We then formalize reward sparsity with observation-anchored reward machines, separating structural sparsity (an automaton reproduces the return without the task-required history) from potential sparsity (the one-step reward misprices local exploratory actions); the resulting vocabulary organizes the three regimes by the retention burden each task exposes. Together, these results show exploration and memory are complements, not substitutes: a bonus induces exposure, and only memory converts exposure into return.

↑