arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化学习中动态奖励塑造的统一框架

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

Fouad Bahrpeyma

arXiv 2608.08158首次发表:更新:

发表机构

Dresden University of Applied Sciences(德累斯顿应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出统一分析框架,区分奖励塑造相关类型,分析12类方法族,明确最优性保证在当代深度强化学习流程中的存续条件,揭示适应率与学习器稳定性间未解决的关系。

AI 中文摘要

稀疏、延迟且信息性弱的奖励仍是高效强化学习的核心障碍。奖励塑造通过补充辅助信号来缓解这些限制,可加速学习,而经典设定中原目标仍为评估准则。成熟理论保证固定塑造信号的安全性:基于势能的奖励塑造在辅助项为时间不变势能的折扣差值时,能保留最优策略。然而在当代强化学习系统中,学习器与可用引导信息在训练中均会演化:价值估计提升、新颖性降低、反馈变化、预测模型优化。自适应奖励机制存在于探索、贝叶斯推理、人在回路学习、自动奖励设计及基于基础模型的方法中。本研究提出统一分析框架,用于比较动态奖励塑造及邻近自适应奖励机制。该框架区分参数修订与状态依赖变化,分离加性塑造、奖励替换及奖励相关引导,并沿时间、信息、理论维度组织现有方法。利用此框架分析了12类方法族,还明确了最优性保证在当代深度强化学习流程、回放缓冲区、自举评论者、奖励归一化中存续的条件,同时揭示了适应率与学习器稳定性间未解决的关系。

英文摘要

Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations by supplementing the task reward with an auxiliary signal that can accelerate learning while, in the classical setting, the original objective remains the evaluation criterion. Established theory guarantees safety for fixed shaping signals: potential-based reward shaping preserves optimal policies when the auxiliary term is the discounted difference of a time-invariant potential. In contemporary reinforcement learning systems, however, both the learner and the information available for guidance evolve during training: value estimates improve, novelty diminishes, feedback shifts, and predictive models are refined. Adaptive reward mechanisms occur across exploration, Bayesian inference, human-in-the-loop learning, automated reward design, and foundation-model-based approaches. This study introduces a unified analytical framework for comparing dynamic reward shaping and neighbouring adaptive reward mechanisms. The proposed framework distinguishes parametric revision from state-dependent variation, separates additive shaping from reward replacement and reward-adjacent guidance, and organises existing methods along temporal, informational, and theoretical dimensions. Using this framework, twelve method families are comparatively analysed. The framework further highlights the conditions under which optimality guarantees survive contemporary deep reinforcement learning pipelines, replay buffers, bootstrapped critics, and reward normalisation, while exposing the unresolved relationship between adaptation rate and learner stability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑