AI 中文总结
本文从游戏算法和基础模型时代的RL两个视角展开研究,前者聚焦多智能体RL中相关概念的相互作用,后者利用生成和基础模型丰富序列决策,开发多种模型并进行相关研究,呈现RL在复杂序列域的统一视图,突出其连接多方面的作用。
AI 中文摘要
强化学习(RL)为明确目标下的序列决策提供框架。经典形式下,RL研究智能体在动态环境中如何行动以最大化长期奖励。在更丰富场景中,问题超出单智能体和固定环境。本文从游戏算法和基础模型时代的RL两个视角研究。第一部分聚焦游戏中的多智能体RL,探讨激励、策略和均衡概念在竞争和一般和环境中的相互作用。第二部分研究基于生成和基础模型的RL,利用先验知识丰富序列决策。本文开发基于扩散的世界模型,研究高效视频生成的RL等,将RL呈现为复杂序列域中目标驱动的适应,突出其在连接决策、环境建模和基础模型能力方面的作用。
英文摘要
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studies how an agent should act to maximise long-term reward in a dynamic environment. In richer settings, the problem extends beyond a single agent and fixed environment: intelligent behavior may require strategic interaction, adaptation to uncertainty, and reasoning over high-dimensional worlds. This thesis studies RL from two perspectives: algorithms in games and RL in the era of foundation models. The first part focuses on multi-agent RL in games. It examines how incentives, policies, and equilibrium concepts interact in competitive and general-sum environments, spanning two-player zero-sum games, large-scale video games, and multi-player settings with general structure. These works investigate learning in multi-agent systems and the behavior of RL methods in interactive environments. The second part studies RL with generative and foundation models, motivated by the idea that prior knowledge can enrich sequential decision making. Pretrained generative models and learned world models serve as representation tools and structured priors for planning, control, and policy optimization. The thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations. It also addresses long-horizon modeling through architectures with memory. Together, these contributions present a unified view of RL as objective-driven adaptation in complex sequential domains. From strategic games to generative world models, the thesis highlights how RL connects decision making, environment modeling, and emerging foundation-model capabilities, offering a broader perspective on the principles underlying intelligent behavior.
CommentsPrinceton University PhD Thesis 2026