arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

微调视觉语言模型以增强AI的空间智能:理解3D和2D旋转

Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations

Uttamasha Monjoree, Wei Yan

arXiv 2610.04206首次发表:更新:

发表机构

Texas A&M University(德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过微调Gemma-4混合专家模型,利用多物体旋转数据集,显著提升了AI在2D和3D旋转检测中的空间智能,并发现线性特征物体有助于角度估计。

AI 中文摘要

空间智能是科学、技术、工程和数学(STEM)、医学、建筑与施工等多个领域的一项基本技能。近期研究表明,视觉语言模型(VLMs)在空间推理方面仍存在局限性,这阻碍了人工智能(AI)执行实际空间任务。利用为训练和评估而开发的多个物体旋转数据集,我们的实验在2D和3D旋转检测方面均展现了令人鼓舞的改进。经过微调的Google DeepMind构建的Gemma-4混合专家(MoE)模型在预测由轴和角度共同定义的旋转方面显著优于微调后的Gemma-4通用模型。微调还大幅提升了无需显式坐标系即可进行的2D表示的角估计。此外,可识别的物体并未提高角度检测的准确性;相反,具有突出线性特征的物体表现出了更好的性能。

英文摘要

Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑