发表机构
Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于人类专家成对轨迹比较的强化学习问题,提出受Bradley-Terry启发的理性模型捕捉不可比性并推断多维奖励函数,分析样本复杂度,评估模型在模拟环境中重建奖励函数及恢复帕累托前沿的能力与稳健性。
AI 中文摘要
在这项工作中,我们研究了来自人类专家提供的成对轨迹比较的强化学习(RL)问题。我们通过形式化一种新设置来泛化基于偏好的RL,在这种设置中,专家还可以将轨迹对标记为不可比,即当两条轨迹都不占主导地位时。我们介绍了学习问题及其解决方案应满足的要求。然后,我们提出了一种受Bradley-Terry启发的新型理性模型,该模型有效地捕捉不可比性并推断多维奖励函数,并研究其性质。当有数据集时,我们对学习模型参数进行了样本复杂度分析。最后,我们评估了我们的模型在模拟环境中重建与专家比较一致的奖励函数以及恢复策略帕累托前沿的能力,以及在不同水平的专家理性下的稳健性分析。
英文摘要
In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither trajectory dominates the other. We introduce the learning problem and the desiderata that its solution should satisfy. Then, we propose a novel Bradley-Terry-inspired rationality model that effectively captures incomparabilities and infers a multi-dimensional reward function, and we study its properties. We provide a sample complexity analysis for learning the model parameters when a dataset is available. Finally, we evaluate our model's ability to reconstruct a reward function that aligns with the expert's comparisons in simulated environments and to recover the Pareto frontier of policies, along with a robustness analysis across varying levels of expert rationality.