发表机构
Texas A&M University(德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过微调Gemma-4混合专家模型,利用多物体旋转数据集,显著提升了AI在2D和3D旋转检测中的空间智能,并发现线性特征物体有助于角度估计。
AI 中文摘要
空间智能是科学、技术、工程和数学(STEM)、医学、建筑与施工等多个领域的一项基本技能。近期研究表明,视觉语言模型(VLMs)在空间推理方面仍存在局限性,这阻碍了人工智能(AI)执行实际空间任务。利用为训练和评估而开发的多个物体旋转数据集,我们的实验在2D和3D旋转检测方面均展现了令人鼓舞的改进。经过微调的Google DeepMind构建的Gemma-4混合专家(MoE)模型在预测由轴和角度共同定义的旋转方面显著优于微调后的Gemma-4通用模型。微调还大幅提升了无需显式坐标系即可进行的2D表示的角估计。此外,可识别的物体并未提高角度检测的准确性;相反,具有突出线性特征的物体表现出了更好的性能。
英文摘要
Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.