arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3387 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3387 篇

2506.02139 2025-12-02 cs.AI 74%

The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning

语言模型的统一认知意识理论:锚定语义、激活阈值与涌现推理

Edward Y. Chang, Zeyneb N. Kaya, Ethan Chang

机构 * Stanford University(斯坦福大学) UIUC(伊利诺伊大学香槟分校)

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

AI总结 该研究提出统一认知意识理论,通过语义锚定解释语言模型如何将预训练能力转化为目标导向行为,并通过实验验证了锚定强度对模型性能的影响。

Comments 21 pages, 7 figure, 4 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16772 2025-10-21 cs.CV cs.AI 74%

Region in Context: Text-condition Image editing with Human-like semantic reasoning

Thuy Phuong Vu, Dinh-Cuong Hoang, Minhhuy Le, Phan Xuan Tan

机构 * Greenwich Vietnam FPT University(越南格林威治FPT大学)

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16643 2025-10-21 cs.CV cs.AI cs.RO 74%

Structured Interfaces for Automated Reasoning with 3D Scene Graphs

Aaron Ray, Jacob Arkin, Harel Biggie, Chuchu Fan, Luca Carlone, Nicholas Roy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14425 2025-08-19 cs.CL 74%

From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning

Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

机构 * Computational Linguistics, Department of Linguistics University of Potsdam(乌特雷赫特大学语言学系计算语言学部) German Research Center for Artificial Intelligence (DFKI), Berlin(德国人工智能研究中心(DFKI)柏林)

专题命中 视觉空间推理 :reasoning(title);分类 cs.CL

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07885 2025-08-12 cs.RO cs.AI cs.CV cs.SY eess.SY 74%

Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning

Shoaib Ahmmad, Zubayer Ahmed Aditto, Md Mehrab Hossain, Noushin Yeasmin, Shorower Hossain

机构 * Department of Mechanical Engineering(机械工程系) Rajshahi University of Engineering and Technology(拉贾沙希工程与技术大学) Department of Industrial and Production Engineering(工业与生产工程系) Shahjalal University of Science and Technology(沙赫jalal科学与技术大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Department of Urban and Regional Planning(城市与区域规划系) Department of Computer Science Engineering(计算机科学与工程系) United International University(联合国际大学)

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01857 2024-12-30 cs.CV cs.LG cs.RO 74%

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang

专题命中 视觉空间推理 :planning(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01652 2024-11-13 cs.RO cs.AI cs.CV 74%

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, Li Fei-Fei

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06237 2024-10-10 cs.RO cs.AI 74%

BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

Rutav Shah, Albert Yu, Yifeng Zhu, Yuke Zhu, Roberto Martín-Martín

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

Comments 7 Figures, 2 Tables, 11 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10103 2024-10-03 cs.RO cs.AI 74%

Reasoning about the Unseen for Efficient Outdoor Object Navigation

Quanting Xie, Tianyi Zhang, Kedi Xu, Matthew Johnson-Roberson, Yonatan Bisk

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

Comments 6 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17080 2024-09-26 cs.CV cs.CL 74%

Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?

Bowen Zhao, Leo Parker Dirac, Paulina Varshavskaya

专题命中 视觉空间推理 :reasoning(title);分类 cs.CL

Comments 13 pages, 4 figures. Code released at https://github.com/groundlight/vlm-visual-demonstrations

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11347 2024-09-18 cs.AI 74%

Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Takanori Ugai, Kensho Hara, Shusaku Egami, Ken Fukuda

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

Comments 5 pages, 1 figure, 1 table, accepted in Embodied AI 2024 Workshop held in conjunction with CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.01397 2020-11-30 cs.RO cs.AI 74%

Guided Navigation from Multiple Viewpoints using Qualitative Spatial Reasoning

Danilo Perico, Paulo E. Santos, Reinaldo Bianchi

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

Comments 26 pages

Journal ref Spatial Cognition and Computation - 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11064 2018-11-28 cs.AI 74%

Combining Deep Learning and Qualitative Spatial Reasoning to Learn Complex Structures from Sparse Examples with Noise

Nikhil Krishnaswamy, Scott Friedman, James Pustejovsky

专题命中 视觉空间推理 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.08565 2018-05-23 cs.LG stat.ML 74%

Global Navigation Using Predictable and Slow Feature Analysis in Multiroom Environments, Path Planning and Other Control Tasks

Stefan Richthofer, Laurenz Wiskott

专题命中 视觉空间推理 :planning(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16952 2025-06-23 cs.CY 73%

Modeling and Visualization Reasoning for Stakeholders in Education and Industry Integration Systems: Research on Structured Synthetic Dialogue Data Generation Based on NIST Standards

Wei Meng

专题命中 视觉空间推理 :reasoning(title,comments)

Comments This paper presents an innovative and rigorous framework for stakeholder modelling in education-industry integration, combining NIST-compliant synthetic data generation with interpretable visual reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05933 2026-06-25 cs.CL cs.AI 版本更新 73%

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

强化学习改善LLM中参数化知识的遍历

Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer abd Niloofar Mireshghallah

机构 * Carnegie Mellon University(卡内基梅隆大学) MIT(麻省理工学院) Stanford University(斯坦福大学)

专题命中 视觉空间推理 :reasoning(abstract);self-correction(abstract);分类 cs.CL、cs.AI

AI总结 本文发现强化学习提升LLM知识回忆能力,原因在于改进了模型对参数化知识层次结构的遍历技能,而非获取新知识。

Comments `

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00095 2026-06-02 cs.CV cs.AI cs.CL cs.RO 73%

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation

弥合2D-3D鸿沟:面向视觉语言导航的分层语义几何地图

Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Bosch Corporate Research(博世企业研究) King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

AI总结 提出分层语义几何地图(HSGM),将3D几何信息转化为VLM可理解的结构化表示,结合VLM高层语义规划与经典路径规划,实现零样本视觉语言导航,在R2R-CE和RxR-CE基准上达到最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20287 2026-05-21 cs.LG cs.AI cs.CV 73%

FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction

FusionCell: 跨注意力融合布局几何与网络列表拓扑以实现标准单元性能预测

Haoyi Zhang, Kairong Guo, Bojie Zhang, Yibo Lin, Runsheng Wang

机构 * School of Integrated Circuits, Peking University, Beijing, China(集成电路学院,北京大学,北京,中国)

专题命中 视觉空间推理 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FusionCell,通过跨注意力机制融合布局几何和网络列表拓扑,以提高标准单元性能预测的准确性,解决了传统方法忽略布局几何导致的耦合和布局依赖效应的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11509 2026-05-13 cs.AI cs.LG cs.MA cs.SY eess.SY 73%

Hierarchical LLM-Driven Control for HAPS-Assisted UAV Networks: Joint Optimization of Flight and Connectivity

分层LLM驱动控制用于HAPS辅助无人机网络:飞行与连接的联合优化

Zijiang Yan, Hao Zhou, Wael Jaafar, Jianhua Pei, Ping Wang, Halim Yanikomeroglu, Hina Tabassum

机构 * Department of Electrical Engineering and Computer Science, York University(约克大学电气工程与计算机科学系) Samsung Research America(三星美国研究院) Department of Software and IT Engineering, École de technologie supérieure (ÉTS), University of Quebec(魁北克大学软件与信息技术工程系,École de technologie supérieure) Non-Terrestrial Networks (Carleton-NTN) Lab and the Department of Systems and Computer Engineering, Carleton University(非地面网络(Carleton-NTN)实验室和系统与计算机工程系,卡尔顿大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在整合陆地和非陆地网络中多无人机系统的联合优化问题,提出基于LLM的分层多速率控制框架,通过高精度仿真平台验证了其在提升运输效率、通信吞吐量和减少碰撞率方面的优势。

Comments Submission for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12597 2026-03-16 cs.LG cs.AI cs.HC cs.MA cs.SE 73%

Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs

Feynman: 一种融合知识的图表生成代理,用于可扩展的视觉设计

Zixin Wen, Yifu Cai, Kyle Lee, Sam Estep, Josh Sunshine, Aarti Singh, Yuejie Chi, Wode Ni

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出Feynman代理,通过知识组件枚举和代码规划生成高质量图表-描述对,构建了可扩展的图表生成流水线,并创建了视觉语言基准Diagramma。

Comments A previous version was submitted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19255 2026-03-06 cs.LG cs.AI 73%

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

VTool-R1: 通过在多模态工具使用上的强化学习使VLMs学会通过图像思考

Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, Chengxiang Zhai, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Michigan Ann Arbor(密歇根大学安娜堡分校) Independent Researcher(独立研究者)

专题命中 视觉空间推理 :reasoning(abstract);self-correction(abstract);分类 cs.AI、cs.LG

AI总结 VTool-R1通过强化学习训练视觉语言模型生成多模态思考链,提升其通过图像进行推理的能力。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01644 2026-02-03 cs.LG cs.AI cs.CV cs.MA cs.RO 73%

From Perception to Action: Spatial AI Agents and World Models

从感知到行动:空间AI代理与世界模型

Gloria Felicia, Nolan Bryant, Handi Putra, Ayaan Gazali, Eliel Lobo, Esteban Rojas

机构 * AtlasPro AI

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种统一的三轴分类法,将代理能力与空间任务联系起来,强调空间定位与符号定位的区别,并指出世界模型对跨尺度安全部署的重要性。

Comments 61 pages, 742 citations, 1 figure, 3 tables. Survey paper on spatial AI agents, embodied AI, graph neural networks, and world models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL 73%

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09775 2026-01-16 cs.LG cs.CL 73%

The Geometry of Thought: Disclosing the Transformer as a Tropical Polynomial Circuit

思维的几何学:揭示Transformer为热带多项式电路

Faruk Alpay, Bilge Senturk

机构 * Bahçeşehir University(巴切希尔大学)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 本研究揭示Transformer在高置信度下通过热带多项式电路实现动态规划,为链式思维提供几何解释。

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02382 2026-01-07 cs.NI cs.AI cs.IR cs.LG 73%

How to Discover Knowledge for FutureG: Contextual RAG and LLM Prompting for O-RAN

如何为未来G发现知识:面向O-RAN的上下文RAG和LLM提示

Nathan Conger, Nathan Scollar, Kemal Davaslioglu, Yalin E. Sagduyu, Sastry Kompella

专题命中 视觉空间推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

AI总结 本文提出上下文RAG方法,通过引导文档检索和上下文增强LLM性能,提升ORAN领域问答的准确性和效率,同时保持低运行时间和碳排放。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23907 2025-12-02 cs.CV cs.AI cs.LG 73%

DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning

DynaStride: 基于MMCoT的动态步长窗多场景描述生成

Eddison Pham, Prisha Priyadarshini, Adrian Maliackel, Kanishk Bandi, Cristian Meo, Kevin Zhu

机构 * Algoverse AI Research(Algoverse人工智能研究)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 DynaStride通过多模态链式思考和动态步长窗算法,实现无需手动分割的多场景教学视频连贯描述生成,提升描述的连贯性和信息量。

Comments 16 pages, 15 figures, 5 Tables, Accepted at NeurIPS 7HVU Workshop, Accepted at AAAI AI4ED Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19433 2025-10-13 cs.CV cs.AI cs.CL 73%

Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System

Lixuan He, Haoyu Dong, Zhenxing Chen, Yangcheng Yu, Jie Feng, Yong Li

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26539 2025-10-01 cs.CV cs.CL cs.LG 73%

Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents

Zhen Yang, Zi-Yi Dou, Di Feng, Forrest Huang, Anh Nguyen, Keen You, Omar Attia, Yuhao Yang, Michael Feng, Haotian Zhang, Ram Ramrakhya, Chao Jia, Jeffrey Nichols, Alexander Toshev, Yinfei Yang, Zhe Gan

机构 * Apple(苹果公司)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11197 2025-09-16 cs.RO cs.AI cs.CL cs.CV 73%

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

Yunheng Wang, Yuetong Fang, Taowen Wang, Yixiao Feng, Yawen Tan, Shuning Zhang, Peiran Liu, Yiding Ji, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhejiang Normal University(浙江师范大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11247 2025-08-19 cs.AI cs.LG cs.RO 73%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 视觉空间推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏