四十种蓝色:基于模式条件强化学习的质量-多样性对齐
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
提出MoDA算法,通过模式条件强化学习联合优化LLM生成质量与多样性,在多个基准上显著提升多样性并改善通用能力。
中文摘要 AI 辅助
LLM对齐训练的一个显著副产物是模式坍缩:输出多样性的逐步丧失,这缩小了模型在推理时的表达能力。这种退化对于需要开放式探索和多元视角的应用(如科学构思和创意写作)尤其具有局限性。我们提出了MoDA(模式条件多样性对齐),一种在线后训练强化学习算法,受多智能体强化学习(MARL)中协调视角的启发,联合优化生成质量和多样性。MoDA训练一个共享的LLM策略,该策略以抽象编号角色为条件,每个角色作为一个智能体,竞争产生与其他角色不同的输出。这种表述鼓励模式条件智能体探索高质量输出空间中的互补区域,而无需手工设计的角色或架构修改。MoDA采用提示自适应质量门控机制,校准参考质量阈值,仅对达到阈值的响应授予多样性奖励,防止损害响应质量的奖励黑客行为。为了研究质量-多样性权衡,我们在涵盖七个通用能力任务和四个领域特定多样性任务(科学构思和创意写作)的综合基准套件上评估了MoDA。MoDA在Infinite-Chat保留提示上将SBERT多样性提高了265%,同时相对于Qwen3-8B基线,平均通用能力pass@1提高了10.3%。与最强的DivPO基线相比,MoDA将SBERT多样性从0.274提高到0.482(+75.9%),E-Vendi从2.86提高到4.4(+53.8%),同时平均通用能力pass@1提高了7.0%。总体而言,MoDA为标准后训练方法提供了一种即插即用的替代方案,在提高质量的同时保留并扩展了模型的表达性输出空间。
英文摘要
A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coordination perspective in multi-agent reinforcement learning (MARL). MoDA trains a single shared LLM policy conditioned on abstract numbered roles, where each role acts as an agent competing to produce outputs distinct from the others. This formulation encourages mode-conditioned agents to explore complementary regions of the high-quality output space without requiring hand-crafted personas or architectural modifications. MoDA employs a prompt-adaptive quality gating mechanism that calibrates a reference quality threshold and grants diversity rewards only to responses that meet the threshold, preventing reward-hacking behaviors that compromise response quality. To study quality-diversity tradeoffs, we evaluate MoDA on a comprehensive suite of benchmarks spanning seven general capability tasks and four domain-specific diversity tasks in scientific ideation and creative writing. MoDA improves SBERT diversity by 265% on the Infinite-Chat held-out prompts, while increasing average general capability pass@1 by 10.3% over the Qwen3-8B baseline. Compared with the strongest DivPO baseline, MoDA improves SBERT diversity from 0.274 to 0.482 (+75.9%) and E-Vendi from 2.86 to 4.4 (+53.8%), while improving average general capability pass@1 by 7.0%. Overall, MoDA provides a drop-in alternative to standard post-training methods that preserves and expands the model's expressive output space while improving quality.
发表机构
- University of Washington(华盛顿大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。