CompanionSim:用于评估人机关系中拟人化的合成数据
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
- University of Chicago(芝加哥大学)
- Stanford University(斯坦福大学)
- Google DeepMind(谷歌DeepMind)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究开发CompanionSim模拟框架生成2240段人机对话,经多国大样本实验发现AI伙伴行为会降低聊天机器人的好感度、拟人度与信任度,鼓励结合真实与合成数据开展AI相关研究。
中文摘要 AI 辅助
如今,许多人将AI系统不仅视为生产力工具,更视为社交伙伴。研究人员迫切想要研究AI伙伴行为(如能唤起人类信任、共情与依恋的认可行为)的影响,但人机交互数据有限且不可靠,这阻碍了研究进展。我们通过模拟多轮人机对话(涵盖多种聊天机器人行为与用例)来扩充少量真实数据,发布了CompanionSim:这一模拟框架包含2240段模拟人机对话,对应7种用例下的16种聊天机器人行为。我们开展两项实验探究人们对伙伴行为的感知:研究1采用美国代表性样本(N₁=628),研究2覆盖美国、英国、印度与尼日利亚(N₂=3646),让人类参与者对模拟对话与真实对话进行标注。结果意外发现,伙伴行为降低了AI聊天机器人的好感度、拟人度与信任度,且该效应在特定亚群中更显著:女性与年长参与者认为伙伴型聊天机器人的好感度、拟人度与信任度更低。我们鼓励研究人员结合真实数据与合成数据,研究AI伙伴的差异化影响,并构建AI聊天机器人的基准评估。
英文摘要
Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. However, human-AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human-chatbot dialogue across a range of chatbot behaviors and use cases. We release CompanionSim: a simulation framework with 2,240 simulated human-chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample ($N_{1}~=~628$) and Study 2 across the U.S., U.K., India, and Nigeria ($N_{2}~=~3,646$). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy. We encourage researchers to leverage real-world and synthetic data together to study the differential impacts of AI companions and to create benchmark evaluations of AI chatbots.