基于正交均衡学习的上下文相关智能体评估
Context-dependent agent evaluation with orthogonal equilibrium learning
浏览论文内容
中文总结 AI 辅助
针对上下文相关智能体评估问题,提出NashEval框架,利用正交均衡学习从离线反馈中稳健识别各上下文中的优胜智能体集合。
中文摘要 AI 辅助
许多应用需要在上下文信息(例如提示、任务或用户群体)下评估智能体。我们研究如何从离线反馈中执行这种上下文相关的智能体评估。现有的用于此目的的基于评分的模型(例如Bradley-Terry模型)强加了传递性偏好排序,当人类判断具有异质性时,这无法反映集体偏好。受社会选择理论的启发,我们将评估构建为两个玩家之间的上下文博弈,每个玩家选择在智能体上的分布作为策略,以获得比对方更大的集体偏好。然后,纳什均衡的支持集定义了特定上下文的获胜者集合。然而,从离线日志中学习特定上下文的均衡是困难的,因为每个上下文仅揭示关于智能体子集的人类反馈,因此,朴素的插值估计器可能有偏。为了解决这些挑战,我们提出了NashEval,一个用于稳健上下文均衡学习的通用框架。NashEval首先构建描述博弈的上下文收益矩阵的去偏估计。然后,NashEval使用定制的正交损失学习上下文到均衡的映射,这避免了为每个上下文单独求解博弈的需要。我们从理论上证明,估计收益矩阵背后的干扰函数的误差仅通过高阶项影响学习到的均衡的风险(即可利用性)。在各种实验中,NashEval提高了均衡学习的稳健性,并一致地识别出跨上下文中表现最佳的智能体集合。
英文摘要
Many applications require to evaluate agents under contextual information (e.g., a prompt, task, or user group). We study how to perform such context-dependent agent evaluation from offline feedback. Existing score-based models for this purpose (e.g., Bradley-Terry) impose a transitive preference ordering, which fails to reflect collective preferences when human judgements are heterogeneous. Inspired by social choice theory, we frame evaluation as a contextual game between two players, each selecting a distribution over agents as the strategy to receive greater collective preference than the other. Then, the support of the Nash equilibrium defines a context-specific set of winners. However, learning context-specific equilibria from offline logs is difficult because each context reveals human feedback on only a subset of agents, and, hence, a naive plug-in estimator can therefore be biased. To address these challenges, we propose NashEval, a general framework for robust contextual equilibrium learning. NashEval first constructs debiased estimates of the contextual payoff matrix that characterizes the game. NashEval then learns the context-to-equilibrium mapping with a tailored orthogonal loss, which avoids the need to solve a separate game for each context. We show theoretically that errors in estimating the nuisance functions underlying the payoff matrix affect the risk of the learned equilibrium (i.e., exploitability) only through higher-order terms. Across various experiments, NashEval improves robustness of equilibrium learning and consistently identifies the set of top-performing agents across contexts.