从垃圾信息中提取有效信号:基于成对比较联合学习奖励与工作者可靠性
Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
浏览论文内容
中文总结 AI 辅助
本研究提出一种基于EM算法的方法,联合学习成对比较场景下的物品奖励与工作者可靠性,通过Polya-Gamma潜变量优化模型,在真实与合成数据集上验证了其对垃圾工作者的强鲁棒性。
中文摘要 AI 辅助
从成对比较中学习的问题已在推荐系统、社会选择乃至近期的大语言模型微调等诸多领域得到广泛研究,该问题的目标是基于物品间的成对比较学习物品的奖励值。在许多场景中,这些比较结果来自Amazon Mechanical Turk、Scale AI等平台的众包工作者;然而,工作者常因领域知识有限或追求收益最大化的垃圾行为而不可靠。本研究旨在探究能否联合学习工作者可靠性(能力)与物品奖励值,为此采用玻尔兹曼理性模型处理成对比较,该模型通过纳入工作者能力扩展了布拉德利-特里-卢斯模型。我们推导了一种基于期望最大化(EM)的学习算法,引入Polya-Gamma潜变量将逻辑似然转化为条件高斯形式,实现可处理的优化并在算法E步得到简化的Q函数;该技术将问题简化为矩阵感知问题,据此为算法建立了理论收敛保证。我们在真实和合成数据集上开展了大量实验,结果表明,与多个基线方法相比,所提算法具有优势,且对垃圾工作者和对抗性工作者均表现出强鲁棒性,凸显了其在实际众包与奖励学习场景中的实用价值。代码和数据已公开于指定网址。
英文摘要
The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to learn item rewards based on pairwise comparisons between them. In many scenarios, these comparisons are elicited from crowdworkers using platforms such as Amazon Mechanical Turk, Scale AI, etc. However, crowdworkers are often unreliable due to limited domain knowledge or revenue-maximizing (spamming) behavior. In this work, our goal is to understand whether worker reliability (competency) can be learned jointly with item rewards. To this end, we adopt the Boltzmann-rational model for pairwise comparisons, which extends the Bradley-Terry-Luce model by incorporating worker competencies. We derive an EM-based algorithm for learning under this model by introducing Polya-Gamma latent variables to transform the logistic likelihood into a conditionally Gaussian form, enabling tractable optimization and leading to a simplified $Q$ function in the E-step of the algorithm. This technique allows us to reduce our formulation to a matrix sensing problem, using which we establish theoretical convergence guarantees for our algorithm. We conduct extensive experiments on real-world and synthetic datasets. These experiments demonstrate the advantages of using our algorithm over several baselines and confirm its strong robustness to both spammers and adversarial workers, highlighting its practical effectiveness in realistic crowdsourcing and reward learning settings. The code and data is publicly available at https://github.com/KaustubhShejole/BoRa_EM.
发表机构
- IIT Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。