主从弱耦合马尔可夫决策过程中的公平策略优化
Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes
浏览论文内容
中文总结 AI 辅助
针对主从弱耦合马尔可夫决策过程,提出公平策略优化框架,利用对称性归约和计数比例深度强化学习,在机器更换及出租车调度中验证了公平与可扩展性。
中文摘要 AI 辅助
我们考虑在建模为主从弱耦合马尔可夫决策过程(M2WCMDP)的序贯决策环境中进行公平资源分配。在该框架中,资源约束将主子马尔可夫决策过程(子MDP)与一组本可独立运行的从属子MDP的动作空间耦合在一起。我们不再使用传统的功利主义(总和)目标,而是优化一类一般的单调、凹、排列不变、归一化的公平函数。对于同质从属子MDP,我们证明在对称性下,该问题可归结为在排列不变策略类上优化平台加平均参与者的功利主义目标,这使得我们能够利用优化基于功利主义目标的高效算法来解决这一公平感知问题。对于更一般的设置,我们引入了一种基于计数比例的深度强化学习方法,并配有一个优先级采样器来生成可行的计数动作。我们框架的通用性意味着所提出的算法和理论保证可迁移到任何具有对称M2WCMDP结构的领域。我们考虑了两个应用:机器更换问题以及纽约市校准数据集上的定价与出租车调度联合控制问题。我们通过全面的实验验证了我们的理论发现,确认了所提方法在实现强公平感知性能的同时保持可扩展性的有效性。
英文摘要
We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textit{permutation-invariant} policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.
发表机构
- GERAD
- HEC Montréal(蒙特利尔高等商学院)
- MILA - Quebec AI Institute(MILA - 魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。