arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用强化微调大型语言模型进行社交推理游戏

Playing social deduction games with reinforcement fine-tuned large language models

Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Kening Zheng, Chiming Duan, Minghua He, Zhaoyang Liu, Bolin Ding, Philip S. Yu, Ying Li

arXiv 2610.04261首次发表:更新:

发表机构

Peking University; Alibaba Group; University of Illinois Chicago(北京大学; 阿里巴巴集团; 伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过社交推理游戏,发现强化微调虽不能直接提升胜负结果,但能有效改善LLM的社交阅读与社交影响力,多智能体社会认知微调可进一步增强其社交能力。

AI 中文摘要

强化微调(RFT)越来越多地应用于大型语言模型(LLMs)与人类及其他智能体交互的场景中。在此,我们利用社交推理游戏来研究RFT如何改变LLMs的社交行为。我们让微调后的和基础LLM智能体参与隐藏角色游戏,这些游戏需要隐藏状态推断、社交阅读和投票引导。我们的结果表明,LLM智能体通过直接优化最终胜负结果并不能可靠地获得社交推理能力,这表明最终游戏结果为社交互动学习提供了稀疏且嘈杂的信号。然而,RFT在改善社交阅读方面特别有效,包括要求智能体从公开讨论中推断隐藏角色、随时间更新信念以及预测其他智能体未来决策的任务。我们进一步表明,RFT还可以改善社交影响力,包括要求智能体引导投票、团队批准和集体决策的任务,尽管这些收益更强烈地依赖于行为特定的奖励和结构化的交互设置。最后,我们表明,通过多智能体社会认知强化微调,LLMs玩社交推理游戏的能力可以进一步提高,该微调在同侧多智能体训练期间结合了社交阅读和社交影响力信号。这些学习到的行为在战略能力、说服力和社交实用性方面也获得了更有利的人类评价。总之,这些结果丰富了我们对RFT如何改变LLMs社交行为的理解,并为机器社会智能的行为学习理论迈出了一步。

英文摘要

Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inference, social reading and vote steering. Our results show that LLM agents do not reliably acquire social-deduction ability by directly optimizing terminal win--loss outcomes, suggesting that final game results provide a sparse and noisy signal for socially interactive learning. However, RFT is particularly effective at improving social reading, including tasks that require agents to infer hidden roles from public discussion, update beliefs over time and predict other agents' future decisions. We further show that RFT can also improve social influence, including tasks that require agents to steer votes, team approvals and collective decisions, although these gains depend more strongly on behaviourally specific rewards and structured interaction settings. Finally, we show that LLMs' ability to play social deduction games can be further improved through multi-agent social-cognitive reinforcement fine-tuning, which combines social-reading and social-influence signals during same-side multi-agent training. These learned behaviours also receive more favourable human evaluations of strategic competence, persuasiveness and social usefulness. Together, these results enrich our understanding of how RFT changes LLMs' social behaviour and provide a step toward a behavioural learning theory for machine social intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑