arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向排名数据的总体级生成建模

Population-Level Generative Modeling for Ranking Data

Zhaoyang Shi

arXiv 2608.08422首次发表:更新:

发表机构

Center for Applied Mathematics, Fudan University(复旦大学应用数学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对排名数据的高维组合与偏好异质性挑战,提出基于潜在偏好单纯形嵌入与流匹配的总体级生成建模框架,在合成与真实数据集上实现更优的总体级排名生成保真度。

AI 中文摘要

排名数据出现在科学与机器学习应用中,包括推荐系统、信息检索、投票、营销以及基于人类反馈的AI偏好排名。现有统计研究主要聚焦于偏好估计、排名聚合和排名预测等推断任务,但从观测总体生成真实的合成排名对于隐私保护数据共享、基准构建、模拟及不确定性量化至关重要。该任务具有挑战性,因为排名是高维组合对象,具有非欧几里得依赖结构,且排名总体常表现出显著的偏好异质性。我们提出一种通过潜在偏好单纯形嵌入实现总体级生成建模的框架:它基于似然排名模型估计低维潜在偏好单纯形,利用流匹配(flow matching)学习潜在偏好的总体分布,并通过拟合的概率排名模型生成新排名。我们证明排名生成可归约为潜在分布学习的最优问题,并推导有限样本生成保证,明确项目数量、排名长度和潜在维度如何影响准确率。在合成与真实数据集上的实验表明,该方法提升了总体级保真度,并提供了偏好异质性的统计可解释表示。

英文摘要

Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggregation, and ranking prediction. However, generating realistic synthetic rankings from an observed population is important for privacy-preserving data sharing, benchmark construction, simulation, and uncertainty quantification. This task is challenging because rankings are high-dimensional combinatorial objects with non-Euclidean dependence structures, while ranking populations often exhibit substantial preference heterogeneity. We propose a framework for population-level generative modeling through a latent preference simplex embedding. It estimates a low-dimensional latent preference simplex through a likelihood-based ranking model, leverages flow matching to learn the population distribution of latent preferences, and generates new rankings through the fitted probabilistic ranking model. We show that ranking generation admits an oracle reduction to latent distribution learning and derive finite-sample generative guarantees that clarify how the number of items, ranking length, and latent dimension affect accuracy. Experiments on synthetic and real datasets demonstrate improved population-level fidelity and provide a statistically interpretable representation of preference heterogeneity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑