发表机构
Seoul National University; Inha University; KAIST; Georgia Institute of Technology(首尔大学; 仁荷大学; 韩国科学技术院; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了3D多模态社交推理游戏测试平台MineAmongUs,提出可配置VLM智能体控制程序ARIA,发现非言语通道是VLM智能体欺骗取胜的更关键因素,为具身VLM智能体对齐研究开辟新方向。
AI 中文摘要
大型语言模型(LLM)和视觉语言模型(VLM)智能体的策略性欺骗已成为AI对齐与安全领域的核心关注点。社交推理游戏(每个玩家持有隐藏角色,通过与他人交流来推断身份)是典型的测试平台,尤其适用于多智能体场景。然而,现有测试平台仅为纯文本形式,且基于单一固定智能体配置,未纳入欺骗分类法视为核心的非言语感知运动通道,还无法明确观察到的行为是源于底层模型还是周边控制程序。本文介绍了MineAmongUs,一个3D多模态《Among Us》沙盒,其中冒名顶替者智能体必须通过联合言语与非言语行动欺骗船员;还提出了ARIA,一种可配置的VLM智能体控制程序,其包含五个认知组件消融轴;以及一种基于欺骗分类法的原子级与弧级标注方案,该方案通过作为评判者的LLM实现大规模操作,达到接近人类的原子标注一致性。实验结果表明,VLM智能体通过联合言语与非言语欺骗追求冒名顶替者胜利,在控制程序消融及跨VLM评估中,非言语通道成为更具决定性的胜利贡献因素。总体而言,本研究为具身VLM智能体的对齐研究开辟了新路径。
英文摘要
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.
CommentsWorkshop on Agent Behavior (WAB) at COLM 2026. Project page: https://junseokim0103.github.io/Lies-We-Can-See/