发表机构
The Hong Kong University of Science and Technology (Guangzhou); Nanyang Technological University; Shanghai Jiao Tong University(香港科技大学(广州); 南洋理工大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有SDG智能体多模态特性缺失的问题,提出首个整合多模态感知生成的CaM-Wolf,经实验验证其游戏性能与交互质量均有提升。
AI 中文摘要
狼人杀等社交推理游戏(SDGs)已成为AI智能体的挑战性测试平台,这类游戏需要推理、欺骗、协作等复杂社交技能。尽管大型语言模型(LLMs)的最新进展推动了SDG智能体的显著进步,但当前方法大多基于文本,忽略了人类社交互动所必需的多模态特性。为弥合这一差距,我们推出CaM-Wolf,这是首个整合多模态感知与生成的SDG智能体。CaM-Wolf处理其他玩家的视频输入,采用通过强化学习训练的因果感知推理器,建立可观测行为与隐藏角色之间的逻辑链,并通过动画化身呈现自身。我们的实验与用户研究表明,CaM-Wolf在智能体游戏玩法上实现了更优性能,并提升了人机交互的质量。这项工作代表着在创建能够参与微妙社交动态、更具人类特性的AI智能体方面取得的重大进展。我们的代码可在此https URL获取。
英文摘要
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.
CommentsAccepted by ACMMM 2026