arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13576cs.LGeess.SP

用于贝叶斯说服的结构化强化学习:在智能交互驾驶中的应用

Structured Reinforcement Learning for Bayesian Persuasion : Application to Intelligent Interactive Driving

Merlin Paul, Anup Aprem

AI总结:

研究智能交互驾驶中主导者引导智能体决策的问题,提出在线结构化强化学习框架,贡献包括为单调智能体提出算法、确定相关条件、提出SQP,数值分析显示该方法在优化行驶奖励上比现有方法成本效率高30%。

AI中文摘要:

交互式驾驶为动态交通管理提供了一种有前景的方法,其中配备实时交通数据的智能引导车辆协调联网车辆的路线选择。本文考虑贝叶斯说服的战略信息揭示框架,主要目标是通过选择性地揭示信息来引导智能体的部分可观测序列决策。然而,智能体的远见性响应使主导者的信号策略设计具有计算挑战性。为此提出了一种在线结构化强化学习框架来合成对有远见的智能体有说服力的计算高效的信号策略。具体贡献包括:为单调智能体提出MAPL算法;确定主导者Q函数超模结构充分条件;确定确保主导者信号策略有说服力的条件;提出Supermodular Q learning for Principal(SQP);数值分析表明所提方法在优化行驶奖励方面比现有方法成本效率高30%。

英文摘要:

Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of harmonising decisions, this paper considers the strategic information revealing framework of Bayesian persuasion. Here, the principal (lead vehicle) aims to guide the agent's (connected vehicle) partially observable sequential decision making towards its own objectives by selectively revealing information, such as real-time traffic ahead, using signals. However, the agent's farsighted response to maximize its long-term reward, renders the principal's signaling strategy design computationally challenging. We propose an online structured reinforcement learning framework to synthesize computationally efficient signaling strategy which is persuasive for a far-sighted agent. The main contributions of the paper are as follows: (i) For a monotonic agent with approximate best response, we propose MAPL, a structured policy learning algorithm for faster online learning, (ii) Identification of sufficient conditions for the supermodular structure of the Q function of the principal for a monotonic agent, (iii) Identification of sufficient conditions to ensure the persuasiveness of the principal's signaling strategy, (iv) Supermodular Q learning for Principal (SQP), which leverages the supermodular structure of principal's action value to synthesize computationally efficient signaling strategy that is persuasive for a monotonic learning agent, (v) Numerical analysis considering a real-time application of Bayesian persuasive driving for lane selection demonstrates that the proposed method is 30% cost efficient for optimising travelling rewards of both the lead and connected vehicle compared to the existing methodologies for signaling strategy design.

↑