arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00175cs.LGcs.AI

用于公平强化学习的推理时策略对齐

Inference-Time Policy Alignment for Fair Reinforcement Learning

Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出乘法策略塑造框架,在不更新预训练RL策略参数的情况下,于推理时提升公平性,同时保留核心任务性能,适用于各类深度RL智能体。

中文摘要 AI 辅助

深度强化学习(RL)智能体通过优化标量奖励函数实现了出色的性能。然而,一旦部署,这些RL智能体的策略往往是固定的,且难以适应新的性能标准。例如,一个训练用于最大化期望累积奖励的智能体可能无法适应之前未知的利益相关者偏好。现有的在RL中实现公平性(一种偏好)的方法通常假设此类偏好是先验已知的,并且需要在面向公平性的指标下对策略进行完整的重新训练。受大语言模型中推理时对齐的启发,我们研究在推理时将预训练的RL策略引导至基于福利的公平性目标的问题,而无需更新基础策略的参数。我们将推理时公平性对齐形式化为策略塑造问题,并提出一种乘法策略塑造框架,该框架使用依赖于动作的福利分数来调整动作概率,因此无需修改基础策略。我们的框架具有通用性,可与任何深度RL智能体兼容。通过在多个领域进行的大量实验,我们证明推理时策略塑造在保持核心任务性能的同时,显著提高了基于福利的公平性目标。

英文摘要

Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For instance, an agent trained to maximize expected cumulative reward may not accommodate previously unknown stakeholder preferences. Existing approaches to achieve fairness, a type of preference, in RL typically assume that such preferences are known a priori and require complete retraining of the policy under a fairness-oriented metric. Inspired by inference-time alignment in large language models, we investigate the problem of steering a pretrained RL policy toward welfare-based fairness objectives at inference time without updating the base policy's parameters. We formalize inference-time fairness alignment as a policy shaping problem and propose a multiplicative policy shaping framework that adjusts action probabilities using action-dependent welfare scores, thus requiring no modification to the base policy. Our framework is general and compatible with any deep RL agent. Through extensive experiments across multiple domains, we demonstrate that inference-time policy shaping substantially improves welfare-based fairness objectives while preserving core task performance.

补充信息

↑