arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失的系数:用于模型个性化的贝叶斯成对合并

The Missing Coefficients: Bayesian Pairwise Merging for Model Personalization

Yaling Shen, Tongtong Wu, Siyuan Yan, Gholamreza Haffari

arXiv 2609.39055首次发表:更新:

发表机构

Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出贝叶斯成对合并(BPM),将用户成对选择转化为模型合并系数,通过后验推断处理反馈模糊性,无需逐用户训练即可实现个性化,并在多项生成任务中验证了有效性。

AI 中文摘要

我们如何根据用户的成对选择来个性化共享专家库?先前的工作可以通过合并奖励专用专家来实现不同的奖励权衡,给定一个权衡权重向量。在实践中,用户更自然地会在输出之间进行选择,而不是指定数值权重。因此,挑战在于将这些选择转化为合并所需的系数,同时考虑反馈有限时的模糊性。我们的关键思想是将未知的奖励权重视为潜在变量:从成对选择和奖励分数差异中推断出它们的后验分布,并直接使用其均值作为合并系数。我们将这一思想实例化为贝叶斯成对合并(BPM),其后验分布还刻画了在给定反馈下哪些奖励权衡仍然合理。我们在放射学摘要、图像描述和故事生成上评估了BPM,涵盖文本到文本和图像到文本生成。在每位模拟用户100条反馈的情况下,BPM相对于均匀合并实现了91.7%、77.1%和64.3%的宏观决定胜率。对于六对模拟用户,每对用户都偏好根据其自身反馈拟合的模型,这一模式也在人类概念验证中观察到。在BPM模型和先验下的模拟中,其温度缩放奖励权重的名义90%区间在仅10次和25次比较时分别实现了88.9%和89.2%的任务平均边际覆盖率。因此,BPM无需为每个用户进行策略训练即可从成对反馈中实现个性化,同时刻画了有限反馈留下的系数模糊性。

英文摘要

How can we personalize a shared expert library from a user's pairwise choices? Prior work can realize different reward trade-offs by merging reward-specialized experts, given a vector of trade-off weights. In practice, users can more naturally choose between outputs than specify numerical weights. The challenge is therefore to turn these choices into the coefficients required for merging, while accounting for ambiguity when feedback is limited. Our key idea is to treat the unknown reward weights as latent variables: infer a posterior over them from pairwise choices and reward-score differences, and use its mean directly as the merge coefficients. We instantiate this idea as Bayesian Pairwise Merging (BPM), whose posterior also characterizes which reward trade-offs remain plausible given the feedback. We evaluate BPM on radiology summarization, image captioning, and story generation, spanning text-to-text and image-to-text generation. With 100 feedback per simulated persona, BPM achieves macro decided win rates of 91.7%, 77.1%, and 64.3% against uniform merge. For six pairs of simulated personas, each prefers the model fitted to its own feedback, a pattern also observed in a human proof-of-concept. In simulations under BPM's model and prior, its nominal 90% intervals for temperature-scaled reward weights achieve task-averaged marginal coverage of 88.9% and 89.2% with only 10 and 25 comparisons, respectively. BPM thus enables personalization from pairwise feedback without per-user policy training, while characterizing the coefficient ambiguity left by limited feedback.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑