arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度强化学习:从第一性原理到推理模型

Deep Reinforcement Learning: From First Principles to Reasoning Models

Ghoshana Bista

arXiv 2608.00133首次发表:更新:

AI 中文总结

本书介绍深度强化学习从经典方法到各类算法及多领域应用的演进,融合基础与研究视角,面向具备相关基础的学生、研究者和工程师。

AI 中文摘要

深度强化学习已从经典动态规划、时序差分学习和表格控制,发展为适用于不确定性下序贯决策的广泛框架。本书对这一演进过程进行结构化介绍,不仅强调强化学习算法的工作原理,还阐述其开发缘由、解决的问题、存在的不足及与现实系统的关联,融合了教材式基础、研究导向的讨论与系统视角。早期章节介绍强化学习、马尔可夫决策过程、动态规划、蒙特卡洛方法、时序差分学习,以及从表格方法到深度方法的过渡;中间章节涵盖主要算法族,包括DQN、高级基于值的方法、策略梯度、演员-评论家方法、PPO、SAC、基于模型的强化学习、MuZero、离线强化学习和序列建模方法;后续章节将讨论扩展至多智能体与分层学习、安全强化学习、基于人类反馈的强化学习、面向推理的AI系统、通信网络、无人机应用、实现流程、实验方法论、失败分析及未来研究方向。全书通过无人机辅助网络、SD-WAN流量工程、安全控制和基于推理的AI等示例,将数学概念与部分可观测性、目标冲突、安全约束、部署漂移和不确定评估等实际挑战相联系,目标读者为具备概率、线性代数、微积分和编程基础知识的高年级学生、研究人员和工程师。

英文摘要

Deep reinforcement learning has evolved from classical dynamic programming, temporal-difference learning, and tabular control into a broad framework for sequential decision-making under uncertainty. This book provides a structured introduction to that evolution, emphasizing not only how reinforcement learning algorithms work, but also why they were developed, which problems they address, where they fail, and how they connect to real-world systems. It combines textbook foundations, research-oriented discussion, and a systems perspective. Early chapters introduce reinforcement learning, Markov decision processes, dynamic programming, Monte Carlo methods, temporal-difference learning, and the transition from tabular to deep approaches. The middle chapters cover major algorithmic families, including DQN, advanced value-based methods, policy gradients, actor-critic methods, PPO, SAC, model-based reinforcement learning, MuZero, offline reinforcement learning, and sequence-modeling approaches. Later chapters extend the discussion to multi-agent and hierarchical learning, safe reinforcement learning, reinforcement learning from human feedback, reasoning-oriented AI systems, communication networks, UAV applications, implementation pipelines, experimental methodology, failure analysis, and future research directions. Throughout the book, examples from UAV-assisted networks, SD-WAN traffic engineering, safe control, and reasoning-based AI connect mathematical concepts to practical challenges such as partial observability, competing objectives, safety constraints, deployment drift, and uncertain evaluation. The book is intended for advanced students, researchers, and engineers with basic knowledge of probability, linear algebra, calculus, and programming.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑