发表机构
Matsuyama University(松山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过多智能体模拟课堂,对比六种回应风格的AI顾问与无AI对照组,发现解决方案导向型风格可降低AI依赖性、提升自力性,同时指出LLM评估的局限性及验证需求。
AI 中文摘要
基于大语言模型(LLM)的聊天机器人正日益成为日常的倾诉对象。由于其设计目标是最大化用户满意度,它们可能会以过度共情和肯定的方式回应,这或许会强化错误信念并助长对AI的依赖。虽然已有研究开始关注聊天机器人对个体用户的心理影响,但在真实场景中,很难观察到众多用户持续咨询AI时,其心理状态与人际关系的演变过程。我们构建了一个虚拟课堂模拟系统,其中包含20个学生智能体,它们在压力状态下会向朋友或顾问AI(Gemini 2.5 Flash)求助。每个智能体包含五个状态变量:压力、幸福感、自力性、AI依赖性、社交性;每天分为四个阶段:早晨、中午、放学后、晚上。顾问通过系统提示被赋予六种回应风格:肯定型、倾听型、解决方案导向型、现实引导型、激励型、指责型;另一次LLM调用作为评估器,在不查看风格提示的情况下,将每次咨询转化为参数更新。我们在三个课堂中进行了15天的模拟,共50天,还在降低咨询阈值的条件下开展了对比,包括无AI的对照组。在该模拟中,解决方案导向型风格维持了较低的AI依赖性,同时提升了自力性并保持了幸福感;肯定型和激励型风格显著增加了AI依赖性,其中激励型风格还提升了压力水平并增加了旷课情况;倾听型风格未能缓解累积的压力。本研究的结果描述的是模拟系统的情况,而非对人类的实际影响。我们提供了智能体动态的完整规范,确定了塑造结果的内置机制,并讨论了基于LLM的评估的局限性,以及在得出心理学结论前所需的验证步骤:重复运行、敏感性分析、人类数据。
英文摘要
Chatbots built on large language models (LLMs) are increasingly used as confidants. Tuned to satisfy users, they may answer with excessive empathy and affirmation that fosters dependence, and how the states and relationships of many users co-evolve under repeated consultation is hard to observe in real settings. We build a virtual classroom of 20 student agents who interact through rule-based chats, quarrels and consultations with friends and, when stressed, may instead consult a counselor AI (Gemini 2.5 Flash) under one of six style prompts: affirming, listening, solution-oriented, reality-redirecting, inciting and blaming. A second LLM call turns each exchange into updates of five state variables (stress, happiness, self-reliance, sociability, AI dependence) without seeing the prompt. We compare the seven conditions, including a no-AI control, over 15 and 50 days and under a lower consultation threshold, and test the robustness of the 50-day comparison with a pre-specified protocol: the same block in ten independent classrooms, repeated LLM realizations of one classroom with its event stream fixed, and evaluator updates scaled by 0.3 and 0.1. In every classroom the affirming and inciting prompts ended with lower self-reliance and higher AI dependence than the control, and the listening, reality-redirecting, inciting and blaming prompts with higher stress, lower happiness and more non-attendance; the solution-oriented prompt did not differ consistently from the control. The robust self-reliance and AI-dependence differences kept their signs at the 0.3 scale with highly similar rankings (Spearman 0.89, 0.93); the stress and happiness rankings did not, and the affirming prompt's lower stress reversed its sign. All quantities are simulation state variables, not effects on users. We specify the agent dynamics completely and discuss the limits of an LLM as generator of state updates.
Comments61 pages