小型语言模型网格中的通信强化学习
Reinforcement Learning of Communication in a Mesh of Small Language Models
浏览论文内容
中文总结 AI 辅助
TalkMesh通过强化学习训练小型语言模型智能体网格学习通信时机与内容,以置信度加权共识提升推理准确率,在GSM8K和MATH-500上显著优于自洽性,并具备对恶意串通的防御能力。
中文摘要 AI 辅助
语言模型通过增加测试时的计算量来提高准确性,但对独立样本进行多数投票会趋于饱和:随着样本数量增加,投票结果收敛于模型最频繁给出的答案。通信可以补充采样所无法提供的能力:一个解决问题的智能体可以将关键步骤传递给其他智能体。我们提出了TalkMesh,一个去中心化的小型语言模型智能体网格,它学习何时通信以及通信什么内容。每个智能体采样一个提案,并用训练好的置信度头对其进行评分。置信度最高的智能体广播一条提示;低于置信度阈值的智能体进行修订,并保留每个得分超过其提案的修订版本。八卦共识在无协调者的情况下近似于按置信度加权的投票。一个谈话策略,使用群体相对策略优化在修订后的正确性变化上进行训练,生成提示和修订。使用三个智能体(它们总共最多生成六个输出),该网格达到了与每个模型进行32次采样多数投票相当的准确率。使用最多8个智能体进行训练,并在32个智能体下进行评估,它将准确率从自洽性下的0.568提高到0.705(Qwen3.5-0.8B,GSM8K),以及从0.492提高到0.722(SmolLM3-3B,MATH-500)。当8个智能体中的4个以伪造的置信度和有毒提示串通给出错误答案时,多数投票准确率降至0.000(Qwen3.5-0.8B,GSM8K)。一个防御性网格,其智能体用自己的置信度头重新评分解决方案,保持了0.507的准确率。在推理、具身协调和交通信号控制任务中,当执行智能体无法观察到所需信息而另一个智能体可以发送该信息时,消息能改善决策。
英文摘要
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500). When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。