基于数据并行吉布斯采样的非参数贝叶斯逆强化学习
Nonparametric Bayesian Inverse Reinforcement Learning with Data-Parallel Gibbs Sampling
浏览论文内容
中文总结 AI 辅助
研究针对多专家示范下逆强化学习问题,采用狄利克雷过程先验的非参数贝叶斯方法,通过数据并行吉布斯采样联合推断潜在奖励类型及奖励,在ObjectWorld评估效果良好,还实现并行加速并刻画权衡。
中文摘要 AI 辅助
逆强化学习从专家示范中恢复奖励函数,但标准公式假设所有示范都来自单个专家。当示范汇集自多个偏好不同的专家时,参数方法恢复的平均奖励不能很好地拟合任何单个专家。我们使用狄利克雷过程先验对奖励函数实现非参数贝叶斯逆强化学习,允许与奖励本身一起推断潜在奖励类型的数量。推理使用一种折叠吉布斯采样器,结合用于聚类分配的中餐厅过程更新和用于奖励权重的 metropolis - hastings 更新,以及软值迭代作为内部规划例程。我们在一个 10x10 的 ObjectWorld 网格上进行评估,有两种和三种真实奖励类型。串行采样器以调整兰德指数 1.000 恢复 K = 2,显著优于最大熵 IRL 基线(ARI = 0.000)。扩展到 K = 3 表明采样器在所有运行中都能正确识别聚类数量;0.48 - 0.58 的分配 ARI 反映了专家类型之间的行为重叠,这种重叠在网格实例中持续存在,表明在 ObjectWorld 上进行可靠的 K = 3 评估需要控制对象放置而不是随机播种。我们还使用 Ray 在 HPC 硬件上跨 CPU 核心并行化采样器,在 8 个工作线程时实现了 4.79 倍的峰值加速,并刻画了状态聚合期间使用的共识合并启发式方法所产生的吞吐量与准确性之间的权衡。代码和容器化环境可在此 https URL 获得。
英文摘要
Inverse Reinforcement Learning recovers reward functions from expert demonstrations, but standard formulations assume that all demonstrations come from a single expert. When demonstrations are pooled from multiple experts with distinct preferences, parametric methods recover an averaged reward that fits no individual expert well. We implement Nonparametric Bayesian Inverse Reinforcement Learning with a Dirichlet Process prior over reward functions, allowing the number of latent reward types to be inferred jointly with the rewards themselves. Inference uses a collapsed Gibbs sampler combining a Chinese Restaurant Process update for cluster assignments with a Metropolis-Hastings update for reward weights, and soft value iteration as the inner planning routine. We evaluate on a 10x10 ObjectWorld grid with two and three ground-truth reward types. The serial sampler recovers K=2 with Adjusted Rand Index of 1.000, substantially outperforming a Maximum Entropy IRL baseline (ARI=0.000). Extension to K=3 shows that the sampler correctly identifies the number of clusters in all runs; assignment ARI of 0.48-0.58 reflects behavioral overlap between expert types that persists across grid instantiations, revealing that reliable K=3 evaluation on ObjectWorld requires controlled object placement rather than random seeding. We further parallelize the sampler across CPU cores using Ray on HPC hardware, achieving a peak speedup of 4.79x at 8 workers, and characterize a throughput-versus-accuracy tradeoff arising from the consensus merge heuristic used during state aggregation. Code and a containerized environment are available at https://github.com/dasashreeya/np_bayes_irl.