arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22170cs.LGcs.AIcs.CL

多重潜在排序能更好地预测语言模型的偏好

Multiple latent orderings better predict language model preferences

Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出语言模型的非传递性偏好源于多个内部一致排序的聚合,引入噪声增强的混合Bradley-Terry模型,在七个模型和四个任务上证明混合排序优于单一效用模型,揭示LLM反映多元偏好,对齐评估需避免平均化不同用户的连贯排序。

中文摘要 AI 辅助

语言模型经常被用于需要做出价值判断和选择的场景中。这些观察到的选择常常表现出非传递性:模型可能偏好项目$A$胜过$B$,偏好$B$胜过$C$,同时却偏好$C$胜过$A$。现有对LLM偏好进行建模的工作将这种不一致性视为围绕单一潜在排序的采样噪声。我们反而提出,非传递性反映了多个潜在、内部一致的排序的聚合。我们首先证明,在任何单调连接函数下,观察到的矛盾都无法由单一排序解释。然后,我们引入一种噪声增强的混合Bradley-Terry(MBT)模型,该模型从重复的成对比较中推断潜在偏好成分。在七个模型和四个任务中,排序的混合体通常比单一效用模型更好地解释结构性矛盾。我们发现,聚合偏好往往隐藏了潜在的偏好异质性。一项关于道德机器困境的案例研究表明,在聚合排序上存在分歧的模型仍可能共享潜在成分。这些结果共同表明,LLM反映了多元偏好。因此,将LLM偏好视为单一函数的对齐和评估流程,可能会对不同用户可能以不同方式认可的连贯排序进行平均化处理,从而带来风险。

英文摘要

Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained by a single ordering under any monotone link function. We then introduce a noise-augmented mixture Bradley-Terry (MBT) model that infers latent preference components from repeated pairwise comparisons. Across seven models and four tasks, a mixture of orderings often explains structural inconsistencies better than single-utility models. We find that aggregate preferences often hide underlying preference heterogeneity. A case study on Moral Machine dilemmas shows that models which disagree on aggregate orderings can still share latent components. Together, these results suggest that LLMs reflect plural preferences. Alignment and evaluation pipelines that treat LLM preferences as a single function, therefore, risk averaging over coherent orderings that different users may endorse differently.

发表机构

  • Vermont Complex Systems Institute(佛蒙特复杂系统研究所)
  • University of Vermont(佛蒙特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑