arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

协变量相关的多元序数偏好联合建模及其与比较模型的联系

Covariate-dependent Joint Modeling of Multivariate Ordinal Preferences and Its Connections with Comparison Models

Yujie Chen, Antik Chakraborty, Anindya Bhadra

arXiv 2610.09070首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出协变量相关的联合连续比率马尔可夫随机场模型,联合建模多元序数偏好,避免数据粗化,并证明其涵盖成对/列表比较模型,且联合建模提升比较效果,同时给出难处理归一化常数下的最大似然推断方法。

AI 中文摘要

多元序数数据连同协变量在从语言模型与人类偏好的对齐到推荐系统等问题中普遍收集。例如,MovieLens等数据集包含人类用户按1-5等级评分的多部电影,以及年龄或性别等人口统计信息。类似地,HelpSteer等数据集收集人类对LLM响应的多个属性(如“有帮助性”或“冗长性”)的序数反馈,协变量取决于LLM提示-响应对。不幸的是,这些数据的标准建模方法(a)单独而非联合地看待属性,(b)通常将数据转换为成对或列表式的胜负比较,以拟合Bradley-Terry和Plackett-Luce等模型。这两种做法都导致实际观测数据的粗化,我们通过一个联合的协变量相关的连续比率马尔可夫随机场模型来解决这一问题。我们还表明,在我们的联合模型的限制条件下可以得到成对或列表式比较模型,并且联合建模改善了比较效果。我们还开发了一种即使在存在难处理的归一化常数的情况下也能进行最大似然推断的程序。

英文摘要

Multivariate ordinal data along with covariates are commonly collected in problems ranging from alignment of language models with human preferences, as well as in recommender systems. For example, data sets such as MovieLens contain several movies rated on a scale 1--5 by human users, along with their demographic information such as age or gender. Similarly, data sets such as HelpSteer collect human feedback on several attributes such as "helpfulness" or "verbosity" of LLM response on an ordinal scale, with covariates depending on the LLM prompt--response pairs. Unfortunately, the standard approaches for modeling these data (a) look at the attributes individually rather than jointly, and (b) often convert the data into pairwise or list-wise win--loss comparisons for fitting models such as Bradley--Terry and Plackett--Luce. Both of these lead to a coarsening of what is actually observed, which we address via a joint covariate-dependent consecutive ratio Markov random field model. We also show pairwise or listwise comparison models are obtained under restrictions of our joint model, and that joint modeling improves comparisons. We also develop a maximum likelihood inference procedure even in the presence of an intractable normalizer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑