SpaRC与SpaRP:面向大型语言模型空间推理能力理解的空间推理表征与路径生成
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
- Hessian Center for AI (hessian.AI)(黑森人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SpaRC框架和SpaRP数据集,系统评估LLMs空间推理能力,发现现有模型表现普遍偏低,但随规模提升,且微调可使F1提高7至32个百分点,闭源模型在拓扑空间推理上仍优于开源模型。
AI中文摘要:
空间推理是生物智能和人工智能的关键组成部分。在本工作中,我们对当前最先进的大型语言模型(LLMs)的空间推理能力进行了全面研究。为支持该研究,我们创建并贡献了一个新颖的空间推理表征(SpaRC)框架和空间推理路径(SpaRP)数据集,以深入理解空间关系与组合以及空间推理链的有用性。我们发现,所有最先进的LLMs在这些数据集上表现不佳——在不同设置下其性能持续偏低。随着模型规模扩大,空间推理能力显著提升。对大型语言模型(如Llama-2-70B)和较小模型(如Llama-2-13B)进行微调,可将其F1分数显著提高7至32个绝对百分点。我们还发现,顶级闭源LLMs在拓扑空间理解与推理方面仍然显著优于其开源对手。
英文摘要:
Spatial reasoning is a crucial component of both biological and artificial intelligence. In this work, we present a comprehensive study of the capability of current state-of-the-art large language models (LLMs) on spatial reasoning. To support our study, we created and contribute a novel Spatial Reasoning Characterization (SpaRC) framework and Spatial Reasoning Paths (SpaRP) datasets, to enable an in-depth understanding of the spatial relations and compositions as well as the usefulness of spatial reasoning chains. We found that all the state-of-the-art LLMs do not perform well on the datasets -- their performances are consistently low across different setups. The spatial reasoning capability improves substantially as model sizes scale up. Finetuning both large language models (e.g., Llama-2-70B) and smaller ones (e.g., Llama-2-13B) can significantly improve their F1-scores by 7--32 absolute points. We also found that the top proprietary LLMs still significantly outperform their open-source counterparts in topological spatial understanding and reasoning.