发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨当每个智能体追求自身奖励时,自我改进的LLM能否通过社会学习提升群体性能,发现复制同伴技能可提高效率但未提升效果。
AI 中文摘要
大型语言模型(LLMs)现在可以通过修改其遵循的指令来改进自身,并且LLM智能体越来越多地被编排在一起解决复杂问题。然而,自我改进方法通常一次优化一个系统,而多智能体框架往往让每个模型朝着共同目标努力。我们提出了一个不同的问题:当每个智能体追求自身奖励时,自我改进的LLM能否充分相互学习,从而提升整个群体的性能?我们将这种能力称为递归社会改进。我们研究那些修改技能文件并选择是否、何时以及向谁复制的群体。独立搜索、向同伴学习以及执行动作共享同一个令牌预算。在受控环境中,已有的社会学习算法能从同伴中获益,但三个LLM却未能如此。它们每令牌获得的奖励低于独立学习者,并且要么探索范围过窄,要么在执行动作前耗尽令牌。随后,我们让模型编写和修改自己的技能。观察同伴改变了它们的改进方式,帮助一个模型更快找到有用技能,另一个模型减少在私有搜索上的开销。然而,两者在相同成本下均未超越独立学习者。技能被复制、修改并传递,因此一项发现可以引发进一步的搜索。然而,这些交流使群体集中于更少的独立发现。综合这些结果,表明LLM可以通过向同伴复制来提高学习效率,但尚不能提升学习效果。
英文摘要
Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent frameworks often have every model work toward a shared goal. We ask a different question. When each agent pursues its own reward, can self-improving LLMs learn from one another well enough to improve the whole population? We call this capability recursive social improvement. We study populations that revise skill files and choose whether, when, and whom to copy from. Independent search, learning from peers, and acting all share one token budget. In controlled environments, established social-learning algorithms benefit from peers, but three LLMs do not. They earn less reward per token than solo learners, and explore too narrowly or run out of tokens before acting. We then let the models write and revise their own skills. Observing peers changes how they improve, helping one model find useful skills sooner and another spend less on private search. Neither, however, outperforms independent learners at the same cost. Skills are copied, revised, and passed on, so one discovery can seed further search. Yet these exchanges concentrate the population around fewer independent discoveries. Together, these results show that LLMs can make learning more efficient by copying from peers, but not yet more effective.