arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35695cs.LGcs.CL

重新思考个性化生成:通过因子化排序模型进行测试时对齐

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

  • University of California, Davis(加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Qiyao Ma, Junshan Zhang, Zhe Zhao

AI总结:

针对个性化生成,提出百万参数MLP排序模型,利用测试时对齐和候选匹配,在九个数据集上以不到0.4%的参数和四个数量级更低的延迟超越十亿参数奖励模型。

AI中文摘要:

将大型语言模型(LLMs)与多样化的用户偏好对齐,从根本上受到标准对齐范式的阻碍,这些范式针对的是单一整体用户。在本工作中,我们首先通过实证研究揭示了在测试时对齐中,个性化生成存在巨大且尚未开发的性能提升空间。我们证明,个性化生成特别适合测试时扩展方法,如最佳N选(BoN),因为它主要可以被视为一个候选匹配问题,而非生成器能力瓶颈。虽然奖励模型原则上可以利用这一提升空间,但它们对个性化校准不佳,且其十亿参数规模使得对大型候选池进行评分成本过高。为克服这一限制,我们提出了一种参数高效的框架,利用百万参数规模的多层感知机(MLP)排序模型。我们的个性化排序模型直接重用基础生成器的内部嵌入,开销极小。通过扩展训练时数据以提供细粒度的个性化偏好,这一百万参数排序模型能准确地对大型候选池进行评分,并能无缝引导生成,降低物化N个候选的成本。在涵盖三种个性化生成设置的九个数据集上的广泛实验表明,我们的个性化排序模型有效利用了所发现的性能提升空间,在每一个数据集上都优于十亿参数规模的通用奖励模型,而其参数不到后者的0.4%,评分延迟低了四个数量级。

英文摘要:

Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.

补充信息

↑