发表机构
Institute of Information Systems, University of Lübeck(信息系统研究所,吕贝克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨如何用视觉语言模型增强人机对话,通过将米斯特拉尔人工智能语言模型与胡椒机器人结合用于人机交互对话,并研究视觉信息对响应时间的影响,发现纳入视觉信息可增添对话背景,且使用欧洲托管的语言模型便于实际应用。
AI 中文摘要
视觉语言模型(VLMs)使机器人能够视觉感知环境以及对话伙伴或协作中人类的动作和特征。对于日常环境中部署的社交机器人及简单自然的使用而言,机器人理解符合人类习惯的情境至关重要。本文介绍了将米斯特拉尔人工智能语言模型与胡椒机器人用于人机交互对话的初步经验,以及研究不同模型中额外视觉信息对响应时间的影响。结果表明,纳入视觉信息能为对话增添背景,响应时间适度增加,使机器人和人类能考虑情境中未言明的因素。此外,使用欧洲托管的语言模型提供了符合欧洲数据保护法规的解决方案,更便于实际应用。
英文摘要
Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.