arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39076cs.AIcs.LG

基于Stackelberg博弈的多语言模型协同对齐

Multi-LLM Collaborative Alignment via Stackelberg Games

Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov

首次发表
浏览论文内容

中文总结 AI 辅助

提出Stackelberg对齐框架,通过EXP3赌博机自适应选择指令,结合声誉加权同伴判断,提升多语言模型协作对齐性能,在12个基准上最高超越基线25%。

中文摘要 AI 辅助

一组语言模型可以通过相互学习彼此的响应来协作并集体改进。这些交互依赖于训练期间使用的指令。现有方法通常均匀地采样指令,尽管随着模型的改进,指令的有用性可能会发生变化:曾经模型响应质量存在差异的指令,之后可能被同等好地解答,而先前困难的指令可能开始提供有用的学习信号。我们提出Stackelberg对齐,一种受博弈论启发的领导者-跟随者框架,将指令选择转变为自适应课程。EXP3赌博机作为领导者,在指令间分配固定采样预算,并使用结合指令难度和响应可区分性的奖励更新其采样分布。语言模型作为跟随者:它们响应选定的指令,评估彼此的响应,并通过DPO或GRPO从产生的偏好信号中学习。该框架使用Elo风格的声誉加权同伴判断和基于声誉的对手匹配,以支持可靠且竞争性的模型交互。在三个异构模型池和12个涵盖科学发现、推理、代码、指令遵循和知识的基准上的实验表明,Stackelberg对齐在三个多样化模型池中实现了最高的宏平均性能,比最强的训练时基线高出最多7.4%,比最佳静态推理基线高出12-25%。分析证实,自适应领导者将决斗集中在信息量最大的指令上,消融研究表明,声誉加权判断和基于声誉的匹配均提高了多语言模型进化的有效性。

英文摘要

A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.

发表机构

  • University of Washington(华盛顿大学)
  • Amazon(亚马逊)
  • University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑