发表机构
NYU; NYU Langone Health(纽约大学; 纽约大学朗格尼医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种无监督算法,从语言模型输出中分离内容与风格,发现九种教师模型中的六种风格,通过重要性加权微调学生模型,在六个数学推理基准上提升Pass@k,实现风格控制并改善推理性能。
AI 中文摘要
语言模型同时学习内容和风格,使得其输出中的风格变化难以识别和控制。我们研究了能否无监督地发现模型响应中的重复风格,并对其进行显式控制。我们设计了一种算法,学习从语言模型的输出中分离内容和风格的表示,并在受控环境中对数学问题验证了其有效性。通过将此方法应用于来自九个不同教师模型的超过10万条已验证轨迹,我们发现了六种重复但分布不均的风格。随后,我们对较小的学生模型进行微调,使其在显式条件下遵循这些风格,并使用重要性加权来平衡语料库中风格贡献的差异。该方法在六个数学推理基准上,相较于对相同数据的标准微调,提高了Pass@$k$指标,证明了我们能有效多样化答案风格。我们确认这也导致了请求风格与实际实现风格之间的强对应关系。我们发现风格影响正确性:解决问题的概率取决于我们施加的风格,不同问题受益于不同风格。总之,我们的结果表明,模型生成数据中的风格变化可以以无监督方式被发现并显式化,为控制和改进推理性能提供了来源。
英文摘要
Language models learn content and style jointly, making stylistic variation in their outputs difficult to identify and control. We study whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. By applying this method to over 100K verified traces from nine distinct teacher models, we discover six recurring yet imbalanced styles. We then fine-tune smaller student models to follow these styles when explicitly conditioned on them, using importance weighting to balance the contribution of the styles represented in the corpus. This approach improves Pass@$k$ over standard fine-tuning on the same data across six math reasoning benchmarks, demonstrating that we can diversify the style of answers effectively. We confirm that this also results in strong correspondence between requested and realized styles. We find that style affects correctness: the probability of solving a problem depends on the style we condition on, and different problems benefit from different styles. In summary, our results show that stylistic variation in model-generated data can be discovered in an unsupervised way, and made explicit, providing a source of both control and improved reasoning performance.
Comments38 pages, 12 figures