发表机构
The IT-University of Copenhagen(哥本哈根信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨短时长开放式社交人机交互中LLM规模与用户感知的关系,发现30B参数模型相比更小模型无显著优势,提示具身社交HRI中模型缩放收益递减,对话流畅性等与参数规模同等重要。
AI 中文摘要
大语言模型(LLMs)正越来越多地被用于驱动具身社交智能体,但更大的模型是否能在短暂的人机交互中提升用户感知仍不明确。本文研究LLM参数规模对机器人界面短时长开放式社交交互的影响。在被试内研究中,19名参与者与由Qwen3-VL模型(参数规模为4B、8B、30B)驱动的机器人面孔进行交互,参与者从感知智能度、自然度、愉悦度和幽默感四个维度对交互进行评价。结果显示,与较小规模变体相比,参与者对30B模型无显著整体偏好,包括在感知自然度或智能度上,30B模型相比4B模型也无显著优势。对于30B模型,AI交互频率与智能度排名之间存在显著关系,这表明更有经验的用户可能对模型能力的差异更敏感。总体而言,研究结果表明,在短时长开放式社交人机交互(HRI)中,模型缩放会产生收益递减效应,对话流畅度、响应速度和社交适宜行为可能与原始参数数量同等重要。
英文摘要
Large Language Models (LLMs) are increasingly used to drive embodied social agents, yet it remains unclear whether larger models improve user perception during brief human-robot encounters. This paper examines the effect of LLM parameter size on short-duration, open-ended social interactions with a robot interface. In a within-subjects study, 19 participants interacted with robot faces driven by Qwen3-VL models at 4B, 8B, and 30B parameters. Participants evaluated the interactions in terms of perceived intelligence, naturalness, enjoyment, and humor. Results showed no significant overall preference for the 30B model over the smaller variants, including no significant advantage over the 4B model in perceived naturalness or intelligence. A significant relationship between AI interaction frequency and intelligence rankings for the 30B model suggests that more experienced users may be more sensitive to differences in model capability. Overall, the findings indicate diminishing returns from model scaling in brief open-ended social HRI, where conversational flow, responsiveness, and socially appropriate behavior potentially matter as much as raw parameter count.
Comments4 pages