arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-30 至 2025-12-30 共收录 5 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 5 篇

2512.22207 2025-12-30 cs.AI 88%

GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks

GamiBench: 通过折纸任务评估多模态大语言模型的空间推理和2D到3D规划能力

Ryan Spencer, Roey Yaari, Ritvik Vemavarapu, Joyce Yang, Steven Ngo, Utkarsh Sharma

专题命中 视觉空间推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI

AI总结 GamiBench通过折纸任务评估多模态大语言模型的空间推理和2D到3D规划能力,引入新指标并揭示模型在复杂折叠任务中的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22183 2025-12-30 cs.CV cs.AI cs.CL 81%

Unbiased Visual Reasoning with Controlled Visual Inputs

通过受控视觉输入实现无偏视觉推理

Zhaonan Li, Shijie Lu, Fei Wang, Jacob Dineen, Xiao Ye, Zhikun Xu, Siyi Liu, Young Min Cho, Bangzheng Li, Daniel Chang, Kenny Nguyen, Qizheng Yang, Muhao Chen, Ben Zhou

机构 * Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学) University of Pennsylvania(宾夕法尼亚大学) University of California, Davis(加州大学戴维斯分校)

专题命中 视觉空间推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 VISTA通过解耦感知与推理,利用受控接口和强化学习训练,提升了视觉推理的鲁棒性和中立性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19683 2025-12-30 cs.CV 78%

From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs

从室内到开放世界:揭示MLLMs中的空间推理差距

Mingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang, Marc Pollefeys, Tong Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) ETH Zürich(苏黎世联邦理工学院) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉空间推理 :reasoning(title,abstract)

AI总结 本文提出了一种大规模基准测试,通过户外数据集揭示MLLMs在空间推理方面的不足,并证明其依赖语言先验而非真实视觉推理。

Comments Project page: https://mingrui-wu.github.io/openbench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18262 2025-12-30 cs.RO cs.AI cs.CV cs.HC cs.LG 62%

ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement

ReSemAct:通过语义结构化和效用细化推进细粒度机器人操作

Chenyu Su, Weiwei Shang, Chen Qian, Fei Zhang, Shuang Cong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 ReSemAct 通过语义结构化和效用细化方法,在细粒度机器人操作中实现更精确的效用目标生成与动态环境适应。

Comments Code and videos: https://github.com/scy-v/ReSemAct and https://resemact.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23024 2025-12-30 cs.CV 50%

With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs

大_context带来大预测能力:通过地-语义场景图进行物体分类

Ciprian Constantinescu, Marius Leordeanu

机构 * National University of Science and Technology(科学与技术国家大学) POLITEHNICA Bucharest(布加勒斯特POLITEHNICA)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 本文提出了一种基于地-语义场景图的上下文感知物体分类框架,通过整合深度估计与全景分割模型,显著提升分类准确率至73.4%,优于传统方法和多模态大语言模型。

Comments This paper is a development of the visual riddle game with Human-AI interaction, entitled "GuessWhat - Riddle Eye with AI", developed by Ciprian Constantinescu (POLItEHNICA Bucharest), Serena Stan (Instituto Cervantes Bucarest) and Marius Leordeanu (POLITEHNICA Bucharest), which was the winner (1st place) of the NeoArt Connect NAC 2025 Scholarship Program

详情

展开后加载摘要…

URL PDF HTML 收藏