arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2402.02330cs.AIcs.CL

增强大语言模型在狼人杀游戏中的推理能力

Enhance Reasoning for Large Language Models in the Game Werewolf

  • Tencent AI Lab(腾讯AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shuang Wu, Liwen Zhu, Tao Yang, Shiwei Xu, Qiang Fu, Yang Wei, Haobo Fu

更新

AI总结:

本文提出将LLM与外部Thinker模块集成的双系统推理框架,通过数据库知识和强化学习训练,在9人狼人杀游戏中提升演绎推理与发言生成能力,并贡献了最大规模社交推理游戏数据集。

AI中文摘要:

本文提出了一个创新框架,将大语言模型(LLMs)与外部Thinker模块集成,以增强基于LLM的智能体的推理能力。与通过提示工程增强LLMs不同,Thinker直接利用数据库中的知识,并采用多种优化技术。该框架形成了一个推理层级,其中LLMs处理直观的系统1任务,如自然语言处理,而Thinker专注于需要复杂逻辑分析和领域特定知识的认知系统2任务。我们的框架通过一个需要双系统推理的9人狼人杀游戏进行展示。我们引入了LLMs与Thinker之间的通信协议,并使用来自18800场人类对局的数据和强化学习来训练Thinker。实验证明了该框架在演绎推理、发言生成和在线游戏评估方面的有效性。此外,我们微调了一个6B参数的LLM,使其在与Thinker集成时超越GPT4。本文还贡献了迄今为止最大的社交推理游戏数据集。

英文摘要:

This paper presents an innovative framework that integrates Large Language Models (LLMs) with an external Thinker module to enhance the reasoning capabilities of LLM-based agents. Unlike augmenting LLMs with prompt engineering, Thinker directly harnesses knowledge from databases and employs various optimization techniques. The framework forms a reasoning hierarchy where LLMs handle intuitive System-1 tasks such as natural language processing, while the Thinker focuses on cognitive System-2 tasks that require complex logical analysis and domain-specific knowledge. Our framework is presented using a 9-player Werewolf game that demands dual-system reasoning. We introduce a communication protocol between LLMs and the Thinker, and train the Thinker using data from 18800 human sessions and reinforcement learning. Experiments demonstrate the framework's effectiveness in deductive reasoning, speech generation, and online game evaluation. Additionally, we fine-tune a 6B LLM to surpass GPT4 when integrated with the Thinker. This paper also contributes the largest dataset for social deduction games to date.

↑