强化学习中用于策略-环境协同设计的环境参数梯度定理
Environment Parameter Gradient Theorem for Co-Design in Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究在环境可变的强化学习中联合优化策略与环境设计参数的问题,通过建立环境参数梯度定理及广义动作值函数,开发无模型算法,在无人机网络设计中证明可联合学习最优放置与路由以降低通信成本。
中文摘要 AI 辅助
传统强化学习关注为固定环境学习控制策略。但在许多工程系统中,环境本身可变,可调整物理或操作参数来塑造智能体经历的转移动态和成本,这促使联合优化策略和环境设计参数。为此,建立了环境参数梯度定理,关键理论工具是广义动作值函数\(Q_{\pi,\xi}(s,a,\zeta)\)。基于此结果开发了无模型算法,在无人机网络设计问题上证明了框架有效性,可联合学习最优无人机放置和通信路由以最小化网络总通信成本。
英文摘要
Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable, i.e., physical or operational parameters can be tuned to shape the system's transition dynamics and costs experienced by the RL agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem --- a formal expression for the gradient of the RL's objective function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{π,ξ}(s,a,ζ)$, which comprises two copies of the environment parameters: $ζ$ governs the cost and transition dynamics at the current state--action pair, while $ξ$ governs the future rollouts. This decoupling yields a tractable closed-form gradient expression and is essential to the theorem's derivation. Building on this result, we develop a model-free algorithm that simultaneously learns the optimal policy and the environment parameters. We demonstrate the efficacy of our framework on a UAV network design problem, where the optimal UAV placement (environment parameters) and communication routes (governed by the policy) are learned jointly to minimize the total communication cost in the network.
发表机构
- Faculty of Mechanical Engineering, Indian Institute of Technology Delhi(机械工程系,印度理工学院德里)
机构由 AI 辅助整理,请以论文原文为准。