小型语言模型能学会谈判吗?强化学习训练卖方受控规模研究
Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
浏览论文内容
中文总结 AI 辅助
本研究通过受控规模实验,使用GRPO训练不同规模的Gemma模型进行谈判,发现调整学习率可显著提升小模型谈判能力,12B模型甚至超越前沿模型。
中文摘要 AI 辅助
大语言模型智能体正开始全面接管客户体验。不久之后,大语言模型可能分别代表公司和客户进行买卖。小模型在规模上更具成本效益,但强化学习能否将它们训练成合格的卖方?我们使用GRPO在程序化效用奖励上训练了四个Gemma 4检查点(有效参数从2.3B到31B),用于双边多议题讨价还价,并在相同的1152场谈判中评估每个分支,对抗两个训练中从未见过的前沿买方。在每种规模使用相同学习率($10^{-6}$)的情况下,强化学习模型相对于其基础模型的增益从2.3B时的$+0.001$上升到31B时的$+0.078$。每种规模仅训练一次,且两个最小的检查点使用不同的架构,因此我们未拟合缩放定律。将学习率提高三倍,在相同或更少的训练步数下,在每种规模上均优于共享学习率,增益从2.3B的$+0.032$到4.5B的$+0.081$。在与两个作为卖方运行的前沿模型的探索性比较中,以三倍学习率训练的12B卖方得分高于两者,尽管其未训练的基础模型得分已与它们相当。以该学习率训练的4.5B卖方与两者均无显著差异,且可适配于一块48 GB GPU。进一步以十倍共享学习率训练的2.3B分支提高了合并得分,但其增益集中在与训练池共享模型族的评估买方上。这些结果表明,在断定小模型无法学会谈判之前,应先调整学习率,并应针对来自多个模型族的买方进行测试。
英文摘要
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.
发表机构
- Fin AI Research
机构由 AI 辅助整理,请以论文原文为准。