arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

聚合可达-规避机会约束下有限规模MDP智能体群体的策略综合

Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints

Jie Fu, Anamika Dubey

arXiv 2610.12028首次发表:更新:

发表机构

University of Florida; Washington State University(佛罗里达大学; 华盛顿州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对有限规模MDP智能体群体的聚合可达-规避机会约束控制问题,通过传播经验密度二阶矩并结合坎泰利不等式,提出基于矩的策略综合方法,在两类场景验证了其性能优于确定性线性规划基线。

AI 中文摘要

考虑具有解耦马尔可夫转移动力学和经验密度反馈的有限规模智能体群体,需满足以下约束:以至少1-δ_r的概率,至少α_r比例的智能体需在某一时刻t^*到达目标区域;同时,在t^*之前的每个时刻,不安全群体比例需以至少1-δ_u的概率保持在β_u以下。然而,标准平均场方法仅在期望意义上强制执行这些约束,无法考虑有限群体规模N下的随机波动。为解决该控制问题,我们通过离散时间李雅普诺夫递推,在平均场轨迹之外同时传播经验密度的二阶矩(方差),并应用坎泰利不等式将机会约束转化为经验密度矩的可处理确定性条件。随后,我们将这些基于矩的代理约束纳入基于梯度的序列凸近似过程,用于密度反馈策略综合。我们进一步引入额外的矩误差界,以构建严格的有限N保证。该方法在网格世界环境和电力系统电动汽车充电聚合问题上进行了评估,并与标准确定性群体水平线性规划基线进行了比较。

英文摘要

Consider a finite population of agents with decoupled Markov transition dynamics and empirical-density feedback, subject to the following constraints: with probability at least $1-δ_r$, at least a fraction $α_r$ of agents must reach a target region at some time $t^*$, while, at each time up to $t^*$, the unsafe population fraction must remain below $β_u$ with probability at least $1-δ_u$. However, standard mean-field methods enforce these constraints only in expectation, which fails to account for stochastic fluctuations at finite fleet size $N$. To address this control problem, we propagate the second-order moment (variance) of the empirical density alongside the mean-field trajectory via a discrete-time Lyapunov recursion, and apply the Cantelli inequality to convert chance constraints into tractable deterministic conditions on the moments of the empirical density. We then incorporate these moment-based surrogate constraints into a gradient-based sequential convex approximation procedure for density-feedback policy synthesis. We further introduce additional moment-error bounds to construct a rigorous finite-$N$ certificate. The method is evaluated on a gridworld environment and a power-system EV-charging aggregation problem and compared with a standard deterministic population-level LP baseline.

Comments8 pages, 2 figures. Submitted to the 2027 American Control Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑