arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-29 至 2025-10-29 共收录 5 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 5 篇

2510.24152 2025-10-29 cs.CV cs.AI 85%

Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning

Aodi Wu, Xubo Luo

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 视觉空间推理 :reasoning(title,abstract);chain-of-thought(abstract);planning(abstract);分类 cs.AI

Comments RoboSense Challenge with IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22118 2025-10-29 cs.CV cs.AI 79%

GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation

Karim Elmaaroufi, Liheng Lai, Justin Svegliato, Yutong Bai, Sanjit A. Seshia, Matei Zaharia

机构 * University of California, Berkeley(加州大学伯克利分校) Models for Embodied and Spatial Harmony(具身与空间和谐模型)

专题命中 视觉空间推理 :reasoning(title,abstract);分类 cs.AI

Comments 22 pages, 3 figures, 3 tables, project page: https://ke7.github.io/graid/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23691 2025-10-29 cs.AI 57%

Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents

Zihao Wang, Xujing Li, Yining Ye, Junjie Fang, Haoming Wang, Longxiang Liu, Shihao Liang, Junting Lu, Zhiyong Wu, Jiazhan Feng, Wanjun Zhong, Zili Li, Yu Wang, Yu Miao, Bo Zhou, Yuanfan Li, Hao Wang, Zhongkai Zhao, Faming Wu, Zhengxuan Jiang, Weihao Tan, Heyuan Yao, Shi Yan, Xiangyang Li, Yitao Liang, Yujia Qin, Guang Shi

机构 * Bytedance Seed(字节跳动种子)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22917 2025-10-29 cs.RO cs.AI 57%

HyPerNav: Hybrid Perception for Object-Oriented Navigation in Unknown Environment

Zecheng Yin, Hao Zhao, Zhen Li

机构 * Future Network of Intelligence Institute(Shenzhen)(智能未来网络研究所(深圳)) Tsinghua University(清华大学) The Chinese University of Hongkong(Shenzhen)(香港大学(深圳))

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 50%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏