arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3389 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3389 篇

2504.07089 2025-06-03 cs.CV cs.CL 57%

OmniCaptioner: One Captioner to Rule Them All

Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao, Qi Qin, Xinyue Li, Le Zhuo, Licheng Wen, Dongyang Liu, Yuewen Cao, Xiangchao Yan, Xin Li, Tianshuo Peng, Shufei Zhang, Botian Shi, Tao Chen, Zhibo Chen, Lei Bai, Peng Gao, Bo Zhang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments More visualizations on Homepage: https://alpha-innovator.github.io/OmniCaptioner-project-page and Official code: https://github.com/Alpha-Innovator/OmniCaptioner

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12821 2025-05-30 cs.CV cs.AI 57%

From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration

Mingyang Song, Xiaoye Qu, Jiawei Zhou, Yu Cheng

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Stony Brook University(石溪大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by CVPR 2025. Project Page: https://vlmlt.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20718 2025-05-29 cs.CV cs.AI 57%

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

Kui Wu, Shuhang Xu, Hao Chen, Churan Wang, Zhoujun Li, Yizhou Wang, Fangwei Zhong

机构 * State Key Laboratory of Complex & Critical Software Environment, Beihang University(北京航空航天大学复杂与关键软件环境国家重点实验室) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) City University of Macau(澳门城市大学) Center on Frontiers of Computing Studies, School of Computer Science, Nat’l Eng. Research Center of Visual Technology, Peking University(北京大学计算机学院前沿计算研究中心、国家工程视觉技术研究中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09546 2025-05-26 cs.RO cs.AI 57%

How Secure Are Large Language Models (LLMs) for Navigation in Urban Environments?

Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang, Geeta Chandra Raju Bethala, Yu-Shen Liu, Mengyu Wang, Anthony Tzes, Yi Fang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13466 2025-05-21 cs.AI 57%

AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data

Vu Dinh Xuan, Hao Vo, David Murphy, Hoang D. Nguyen

机构 * University of Information Technology, VNU–HCM(越南胡志明市信息技术大学) University College Cork(科尔克大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03368 2025-05-13 cs.LG 57%

Geospatial Mechanistic Interpretability of Large Language Models

Stef De Sabbata, Stefano Mizzaro, Kevin Roitero

机构 * University of Leicester, UK(利兹大学) University of Udine, Italy(乌迪内大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Figures 2 and 3: fixed issue with min boundary in colorbar

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05485 2025-05-13 cs.RO cs.AI cs.CV 57%

HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Yi Li, Yuquan Deng, Jesse Zhang, Joel Jang, Marius Memmel, Raymond Yu, Caelan Reed Garrett, Fabio Ramos, Dieter Fox, Anqi Li, Abhishek Gupta, Ankit Goyal

机构 * NVIDIA University of Washington(华盛顿大学) University of Southern California(南加州大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments update related work and results on VQA benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04718 2025-05-09 cs.CV cs.LG 57%

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

Divyansh Srivastava, Xiang Zhang, He Wen, Chenru Wen, Zhuowen Tu

机构 * UC San Diego(加州大学圣地亚哥分校) Tsingua University(清华大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07113 2025-05-07 cs.CV cs.AI 57%

Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph

Sergey Linok, Tatiana Zemskova, Svetlana Ladanova, Roman Titkov, Dmitry Yudin, Maxim Monastyrny, Aleksei Valenkov

机构 * Center for Cognitive Modeling, Moscow Institute of Physics and Technology(认知建模中心,莫斯科物理技术学院) AIRI Sberbank of Russia, Robotics Center(俄罗斯储蓄银行机器人中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07937 2025-05-07 cs.RO cs.AI 57%

Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

Bangguo Yu, Qihao Yuan, Kailai Li, Hamidreza Kasaei, Ming Cao

机构 * Faculty of Science and Engineering, University of Groningen(工程学院,格罗宁根大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02829 2025-05-06 cs.AI 57%

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta, Trevor Darrell, David M. Chan

机构 * Department of Electrical Engineering and Computer Sciences, University of California-Berkeley, Berkeley, CA, USA(电气工程与计算机科学系,加州大学伯克利分校)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 28 pages, 10 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11564 2025-05-01 cs.RO cs.CL cs.CV cs.HC 57%

Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding

Yaoxian Song, Penglei Sun, Piaopiao Jin, Yi Ren, Yu Zheng, Zhixu Li, Xiaowen Chu, Yue Zhang, Tiefeng Li, Jason Gu

机构 * Center for X-Mechanics, School of Aeronautics and Astronautics, Zhejiang University(浙江大学航空航天学院X力学中心) Information Hub, Data Science and Analytics Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心、数据科学与分析方向) Advanced Manufacturing Lab, Huawei Technologies(华为技术先进制造实验室) Research Institute, UBTECH Robotics Inc.(UBTECH机器人研究院) School of Information and School of Smart Governance, Renmin University of China(中国人民大学信息学院和智能治理学院) School of Engineering, Westlake University and the Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖大学工程学院和西湖先进科技研究院) Department of Electrical and Computer Engineering, Dalhousie University(达尔豪斯大学电气与计算机工程系)

专题命中 视觉空间推理 :planning(abstract);分类 cs.CL

Comments 15 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11419 2025-04-29 cs.AI cs.NE 57%

Embodied World Models Emerge from Navigational Task in Open-Ended Environments

Li Jin, Liu Jia

机构 * Tsinghua Laboratory of Brain and Intelligence(清华大学脑科学与智能实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Research on explainable meta-reinforcement learning AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12533 2025-04-25 cs.RO cs.AI 57%

To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions

Daniel Tanneberg, Felix Ocker, Stephan Hasler, Joerg Deigmoeller, Anna Belardinelli, Chao Wang, Heiko Wersing, Bernhard Sendhoff, Michael Gienger

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16791 2025-04-22 cs.HC cs.AI 57%

"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation

Zijian Ding, Michelle Brachman, Joel Chan, Werner Geyer

机构 * College of Information, University of Maryland USA(信息学院,马里兰大学美国分校) IBM Research USA(IBM美国研究)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13942 2025-04-22 cs.HC cs.AI cs.ET 57%

Intelligence of Things: A Spatial Context-Aware Control System for Smart Devices

Sukanth Kalivarathan, Muhmmad Abrar Raja Mohamed, Aswathy Ravikumar, S Harini

机构 * School of Computer Science and Engineering, Vellore Institute of Technology(计算机科学与工程学院,维杰里学院技术学院) Sustainable Living Labs(可持续生活实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 16 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12817 2025-04-18 cs.RO cs.AI cs.CV 57%

Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks

Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Workshop "Advancing Automated Driving in Highly Interactive Scenarios through Behavior Prediction, Trustworthy AI, and Remote Operations" @ 36th IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17385 2025-04-18 cs.CL cs.CV 57%

Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Zheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Chai, Ziqiao Ma

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments Accepted to ICLR 2025 (Oral) | Project page: https://spatial-comfort.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10784 2025-04-16 cs.RO cs.AI 57%

ATLASv2: LLM-Guided Adaptive Landmark Acquisition and Navigation on the Edge

Mikolaj Walczak, Uttej Kallakuri, Tinoosh Mohsenin

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10669 2025-04-16 cs.CV cs.LG 57%

Perturbed State Space Feature Encoders for Optical Flow with Event Cameras

Gokul Raju Govinda Raju, Nikola Zubić, Marco Cannici, Davide Scaramuzza

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments 10 pages, 4 figures, 4 tables. Equal contribution by Gokul Raju Govinda Raju and Nikola Zubić

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07615 2025-04-15 cs.CV cs.CL 57%

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, Ruochen Xu, Tiancheng Zhao

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments 11 pages, fix some minor typos in the previous version

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00569 2025-04-15 cs.CV cs.LG 57%

Probing Visual Language Priors in VLMs

Tiange Luo, Ang Cao, Gunhee Lee, Justin Johnson, Honglak Lee

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Project Page: https://vilp-team.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07198 2025-04-11 cs.CV cs.AI cs.HC 57%

Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning

Ashutosh Chaubey, Xulang Guan, Mohammad Soleymani

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Project Page: https://face-llava.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04303 2025-04-08 cs.CL 57%

Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Yue Dai, Soyeon Caren Han, Wei Liu

专题命中 视觉空间推理 :CoT(abstract);分类 cs.CL

Comments Accepted at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01708 2025-04-03 cs.RO cs.HC cs.LG 57%

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication

Petr Vanc, Karla Stepanova

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12836 2025-04-03 cs.GR cs.AI cs.CV cs.HC 57%

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

Kaizhi Zheng, Xiaotong Chen, Xuehai He, Jing Gu, Linjie Li, Zhengyuan Yang, Kevin Lin, Jianfeng Wang, Lijuan Wang, Xin Eric Wang

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19680 2025-03-28 cs.CV cs.AI 57%

M-LLM Based Video Frame Selection for Efficient Video Understanding

Kai Hu, Feng Gao, Xiaohan Nie, Peng Zhou, Son Tran, Tal Neiman, Lingyun Wang, Mubarak Shah, Raffay Hamid, Bing Yin, Trishul Chilimbi

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20241 2025-03-27 cs.RO cs.AI 57%

LGR: LLM-Guided Ranking of Frontiers for Object Goal Navigation

Mitsuaki Uno, Kanji Tanaka, Daiki Iwata, Yudai Noda, Shoya Miyazaki, Kouki Terashima

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 10 pages, 11 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18476 2025-03-26 cs.CV cs.CL 57%

Global-Local Tree Search in VLMs for 3D Indoor Scene Generation

Wei Deng, Mengshi Qi, Huadong Ma

专题命中 视觉空间推理 :planning(abstract);分类 cs.CL

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18730 2025-03-25 cs.CL 57%

Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving

Hongkuan Zhou, Stefan Schmid, Yicong Li, Lavdim Halilaj, Xiangtong Yao, Wei cao

专题命中 视觉空间推理 :planning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏