arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体AI结合结构化思维链以增强AI的空间智能:旋转的可视化与推理

Agentic AI with Structured CoT for Enhancing AI's Spatial Intelligence: Visualization and Reasoning of Rotation

Uttamasha Monjoree, Wei Yan

arXiv 2610.04188首次发表:更新:

发表机构

Texas A&M University(德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过结构化思维链结合上下文信息,显著提升了GPT-5.6模型在三维旋转任务上的空间推理性能,揭示了视觉、文本与推理信息整合对智能体AI空间智能的关键作用。

AI 中文摘要

近期研究表明,具备语言和视觉能力的人工智能(AI)在空间推理方面仍存在局限性。本文研究了先进生成式AI理解三维空间中物体旋转的空间能力,利用AI的图像处理和语言处理特性。我们训练并检验了一个生成式智能体AI模型(GPT-5.6)的空间智能,使其基于修订版普渡空间可视化测试:旋转可视化(修订版PSVT:R)中的旋转图来理解空间旋转过程。我们通过在修订版PSVT:R上叠加额外的图形和上下文特征,评估了不同思维链(CoT)推理策略对模型性能的影响。结果表明,结构化CoT推理在两种数据集(PSVT:R和带坐标系的PSVT:R)上均提升了基础GPT-5.6模型的空间推理性能。我们使用了三种CoT方法:(1)结构化CoT,(2)少样本结构化CoT,以及带自优化提示的结构化CoT。本研究中评估的三种CoT方法在性能上没有显著差异。结果显示,将结构化CoT推理与相关上下文信息相结合,能显著提升视觉语言模型(VLM)在三维旋转任务上的性能,展示了智能体AI在更有效空间推理方面的潜力。然而,当移除上下文信息时,单独的结构化CoT推理仅带来有限的改进,模型在理解空间变换方面仍表现出显著困难。这些发现表明,在未来的智能体AI系统中,有效的空间推理依赖于视觉、文本和基于推理的信息的整合,以实现空间智能。

英文摘要

Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In this paper, we have studied the spatial capabilities of advanced generative AI to understand the rotations of objects in 3D space, utilizing AI's image processing and language processing features. We trained and examined the spatial intelligence of a generative Agentic AI model (GPT-5.6) to understand the spatial rotation process with rotation diagrams based on the revised Purdue Spatial Visualization Test: Visualization of Rotations (Revised PSVT:R). We improvised the Revised PSVT:R by superimposing additional graphical and contextual features to evaluate how different Chain-of-Thought (CoT) reasoning strategies influence model performance. The results indicate that structured CoT reasoning improves the spatial reasoning performance of the base GPT-5.6 model in both datasets (PSVT:R and PSVT:R with coordinate system). We used three CoT approaches - (1) Structured CoT, (2) few-shot Structured CoT, and Structured CoT with Self-optimized Prompt. The three CoT approaches evaluated in this study showed no significant performance difference. Results showed that combining structured CoT reasoning with relevant contextual information leads to considerable improvements in VLM performance on 3D rotation tasks, demonstrating the potential of agentic AI for more effective spatial reasoning. However, when contextual information is removed, structured CoT reasoning alone provides limited improvement, and the models continue to exhibit notable difficulties in understanding spatial transformations. These findings suggest that effective spatial reasoning in VLMs relies on the integration of visual, textual, and reasoning-based information in future agentic AI systems for spatial intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑