Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
在家庭环境中评估多模态大语言模型的日常复合任务
机构 * State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence(通用人工智能国家重点实验室、北京通用人工智能研究院) ; School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Key Laboratory of Machine Perception (Ministry of Education), Peking University(心理与认知科学学院及北京行为与心理健康重点实验室、机器感知重点实验室(教育部)) ; School of Intelligence Science and Technology, Peking University(智能科学与技术学院)
专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.AI
AI总结 本研究在家庭环境中评估了多模态大语言模型在日常复合任务中的表现,发现其在物体理解、空间智能和社会活动领域存在显著差距,为具身MLLMs的发展提供了初步评估框架。