反事实沙普利信用分配
Counterfactual Shapley Credit Assignment
AI总结:
研究信用分配问题,提出基于因果理论的反事实沙普利信用分配框架,通过反事实沙普利值分配功劳责任,推导有效估计器实现$\phi$-PPO方法,结合优先轨迹重放,在复杂环境中有卓越样本效率,能精确对齐奖励真实原因。
AI中文摘要:
信用分配问题(CAP)对于开发高效且可解释的强化学习(RL)智能体至关重要。现有框架常无法在智能体策略(技能)和环境随机性(运气)之间正确归因。一种有原则的CAP方法必须从虚假关联和环境随机性中分离出观察结果的真正因果驱动因素。我们引入反事实沙普利信用分配,这是一个基于因果理论的新框架,通过反事实沙普利值($\phi$值)来分配功劳和责任。通过重新分配环境奖励,$\phi$值在稀疏因果关系、高随机性和延迟奖励这三个关键维度上增强了时间信用分配,同时保留最优策略。我们推导了一个能有效计算$\phi$值的一致估计器,实现了一类新的策略梯度方法$\phi$-PPO,并结合优先轨迹重放(PTR)。实证结果表明,在先前最先进方法无法收敛的具有挑战性的环境中,$\phi$值与任务奖励的真实原因精确对齐,且具有卓越的样本效率。
英文摘要:
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to CAP must isolate the true causal drivers of observed outcomes from spurious correlations and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel framework grounded in causal theory that attributes credit and blame via the Counterfactual Shapley Value ($ϕ$-value). By redistributing environmental rewards, $ϕ$-values enhance temporal credit assignment across three critical dimensions: sparse causality, high stochasticity, and delayed rewards, all while preserving the optimal policy. We derive a consistent estimator that computes $ϕ$-values efficiently, enabling a new class of policy gradient methods, $ϕ$-PPO, combined with Prioritized Trajectory Replay (PTR). Empirical results demonstrate that $ϕ$-values align precisely to the ground truth causes of task rewards with superior sample efficiency in challenging environments where prior state-of-the-art methods fail to converge.