arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45335 信号源:cs.CL, cs.AI, cs.LG

1. 规划推理 11088 篇

2510.18095 2026-02-13 cs.AI cs.CL 88%

SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning

SMaRT: 选择、混合与再发明 -- 一种用于大语言模型驱动推理与规划的策略融合框架

Nikhil Verma, Manasa Bharadwaj, Wonjun Jang, Harmanpreet Singh, Yixiao Wang, Homa Fashandi, Chul Lee

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

AI总结 SMaRT框架通过融合多样化推理策略,提升大语言模型在推理与规划任务中的性能和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14406 2026-02-11 cs.AI cs.CL 88%

IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning

IMAGINE:将多智能体系统整合到一个模型中以实现复杂推理和规划

Xikai Zhang, Bo Wang, Likang Xiao, Yongzhi Li, Quan Chen, Wenjun Wu, Liu Liu

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Kuaishou Technology(快手科技)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

AI总结 IMAGINE通过整合多智能体系统的推理与规划能力,实现高效复杂推理和规划,实验显示其在TravelPlanner任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04326 2026-02-05 cs.AI cs.CL cs.MA 88%

From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents

从假设到行动:将大语言模型推理转化为具有不确定性的代理规划

SeungWon Seo, SooBin Lim, SeongRae Noh, Haneul Kim, HyeongYeop Kang

机构 * Department of Computer Science and Engineering, Korea University(计算机科学与工程系,韩国大学)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

AI总结 PCE框架通过结构化决策树将LLM推理中的假设转化为可靠策略,提升多智能体环境下的规划效率和任务完成率。

Comments 31 pages, 10 figures, Accepted ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14456 2026-01-22 cs.AI cs.LG 88%

On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL

在LLM规划中的泛化差距:测试与验证者奖励强化学习

Valerio Belcamino, Nicholas Attolino, Alessio Capitanelli, Fulvio Mastrogiovanni

机构 * Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genoa(信息学、生物工程、机器人学和系统工程系,热那亚大学) AIKO S.r.l.(AIKO公司)

专题命中 规划推理 :planning(title,abstract);verifier(title,abstract);分类 cs.AI、cs.LG

AI总结 本研究发现微调LLM在规划任务中存在显著的泛化差距,通过三种诊断干预揭示模型依赖领域特定模式而非可转移能力。

Comments 9 pages, 4 figures, 3 tables, 2 pages of supplementary materials. Submitted to a conference implementing a double-blind review process

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05745 2025-12-04 cs.AI cs.LG 88%

SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models

SPRINT: 使推理模型能够实现交错规划与并行执行

Emil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza, Azalia Mirhoseini, Amin Saberi

机构 * Stanford University(斯坦福大学) Microsoft(微软) Google(谷歌)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG

AI总结 SPRINT通过动态识别并利用并行化机会,使推理模型在复杂任务中提升效率,减少序列token生成量。

Comments Published at NeurIPS 2025. Emil Biju, Shayan Talaei, and Zhemin Huang contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22315 2025-11-12 cs.AI cs.CL 88%

PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning

Hieu Tran, Zonghai Yao, Nguyen Luong Tran, Zhichao Yang, Feiyun Ouyang, Shuo Han, Razieh Rahimi, Hong Yu

机构 * Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA Bedford Health Care 医疗组织与实施研究中心) Manning College of Information and Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校信息与计算机科学学院) Miner School of Computer and Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校 Miner 信息与计算机科学学院)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments Proceedings of AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02652 2025-11-03 cs.AI cs.CL cs.IR 88%

HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search

Jiajie Jin, Xiaoxi Li, Guanting Dong, Yuyao Zhang, Yutao Zhu, Yang Zhao, Hongjin Qian, Zhicheng Dou

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23870 2025-10-29 cs.CL cs.AI 88%

OraPlan-SQL: A Planning-Centric Framework for Complex Bilingual NL2SQL Reasoning

Marianne Menglin Liu, Sai Ashish Somayajula, Syed Fahad Allam Shah, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20096 2025-10-14 cs.CL cs.AI 88%

MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning

Thang Nguyen, Peter Chin, Yu-Wing Tai

机构 * Dartmouth College(达特茅斯学院)

专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06475 2025-10-09 cs.AI cs.CL 88%

PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles

Yitao Long, Yuru Jiang, Hongjun Liu, Yilun Zhao, Jingchen Sun, Yiqiu Shen, Chen Zhao, Arman Cohan, Dennis Shasha

机构 * New York University(纽约大学) Zhejiang University(浙江大学) Yale University(耶鲁大学) University at Buffalo, SUNY(布法罗大学) NYU Grossman School of Medicine(纽约大学医学院)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25052 2025-09-30 cs.AI cs.LG 88%

Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning

Sai Wang, Yu Wu, Zhongwen Xu

机构 * Tencent(腾讯) Wuhan University(武汉大学)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13070 2025-08-19 cs.CL cs.AI 88%

Reinforced Context Order Recovery for Adaptive Reasoning and Planning

Long Ma, Fangwei Zhong, Yizhou Wang

机构 * Peking University(北京大学) Beijing Normal University(北京师范大学)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22050 2025-07-15 cs.AI cs.LG 88%

Reinforced Reasoning for Embodied Planning

Di Wu, Jiaxin Fan, Junzhe Zang, Guanbo Wang, Wei Yin, Wenhao Li, Bo Jin

机构 * Tongji University(同济大学) Tsinghua University(清华大学) Bank of Communications(中国银行)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01475 2025-06-03 cs.AI cs.CL 88%

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization

Zouying Cao, Runze Wang, Yifei Yang, Xinbei Ma, Xiaoyong Zhu, Bo Zheng, Hai Zhao

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments 20 pages, 12 figures, 14 tables, ACL'25 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14157 2025-02-19 cs.CL cs.LG 88%

Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning

Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, Lingpeng Kong

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17195 2024-10-29 cs.AI cs.CL 88%

Non-myopic Generation of Language Models for Reasoning and Planning

Chang Ma, Haiteng Zhao, Junlei Zhang, Junxian He, Lingpeng Kong

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20007 2024-10-29 cs.AI cs.CL 88%

Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models

Danqing Wang, Zhuorui Ye, Fei Fang, Lei Li

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18963 2024-10-25 cs.AI cs.CL 88%

OSCAR: Operating System Control via State-Aware Reasoning and Re-Planning

Xiaoqiang Wang, Bang Liu

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19479 2024-10-01 cs.LG cs.AI 88%

Spatial Reasoning and Planning for Deep Embodied Agents

Shu Ishida

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG

Comments DPhil Thesis - Engineering Science, University of Oxford. Original copy available at https://ora.ox.ac.uk/objects/uuid:19489c19-dc5a-464a-831d-bbf887687c41

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15278 2024-05-28 cs.RO cs.AI cs.CV cs.LG 88%

Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models

Yixuan Huang, Jialin Yuan, Chanho Kim, Pupul Pradhan, Bryan Chen, Li Fuxin, Tucker Hermans

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG

Comments Presented at IEEE Conference on Robotics and Automation (ICRA) 2024. Website: https://sites.google.com/view/rdmemory

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11584 2024-04-18 cs.AI cs.CL 88%

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey

Tula Masterman, Sandi Besen, Mason Sawtell, Alex Chao

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments 13 pages,6 figures,38 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10498 2023-11-28 cs.CL cs.AI 88%

PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, Subbarao Kambhampati

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2023 Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13100 2026-03-16 cs.RO cs.AI 88%

Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences

评估VLMs在机器人运动中的空间推理能力:迈向具有运动偏好的机器人规划的一步

Wenxi Wu, Jingjing Zhang, Martim Brandão

机构 * King’s College London(伦敦国王学院) University College London(伦敦大学学院)

专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI

AI总结 本文评估了四种先进VLMs在机器人运动中的空间推理能力,探讨了运动偏好(如物体接近性和路径风格)的处理,并分析了准确率与计算成本的权衡。

Comments Accepted to the First Workshop on Efficient Spatial Reasoning at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21543 2026-06-24 cs.RO 版本更新 88%

Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation

Self-CriTeach: LLM 自我教学与自我批评用于通过自动领域生成提升机器人规划

Jinbang Huang, Zhiyuan Li, Yuanzhao Hu, Zhanguang Zhang, Mark Coates, Xingyue Quan, Yingxue Zhang

机构 * Huawei Noah's Ark Lab(华为诺亚实验室) University of Toronto(多伦多大学) University of British Columbia(不列颠哥伦比亚大学) McGill University(麦吉尔大学)

专题命中 规划推理 :planning(title,abstract);CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出Self-CriTeach框架,通过LLM自动生成符号规划领域,用于自我教学生成规划问题-计划对及自我批评生成结构化奖励信号,提升机器人规划性能与泛化能力。

Comments International Conference on Machine Learning (ICML) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17937 2026-06-17 cs.RO 新提交 88%

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

ThinkingVLA:用于机器人操作的交叉视觉与语言推理

Tianyi Lu, Hui Zhang, Zijie Diao, Junke Wang, Shengqi Xu, Xingyao Lin, Guojin Zhong, Ziyi Ye, Peng Wang, Zuxuan Wu, Yu-Gang Jiang

机构 * Fudan University(复旦大学)

专题命中 规划推理 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);planning(abstract)

AI总结 提出ThinkingVLA,通过统一的多Transformer架构实现前向与逆向推理交织,显著提升长时域操作任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09551 2026-05-27 cs.CV 88%

GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision

GeoSolver: 利用细粒度过程监督扩展遥感中的测试时推理

Lang Sun, Ronghao Fu, Zhuoran Duan, Haoran Liu, Xueyan Liu, Bo Yang

机构 * College of Computer Science and Technology(计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University(教育部符号计算与知识工程重点实验室)

专题命中 规划推理 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);verifier(abstract)

AI总结 提出GeoSolver框架,通过构建大规模过程监督数据集Geo-PRM-2M和训练过程奖励模型GeoPRM,结合过程感知树GRPO强化学习算法,实现遥感中可验证的逐步推理,在多个基准上达到最优性能并支持测试时扩展。

Comments Code: https://github.com/yourname/GeoSolver

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23460 2026-04-28 cs.AI cs.CL cs.LG 88%

Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models

隐秘动机:在连续思维模型中检测不一致的推理

Sharan Ramjee

机构 * Stanford University(斯坦福大学)

专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)

AI总结 本文研究了连续思维模型中如何检测不一致的推理,提出MoralChain基准测试,并通过双触发方法训练模型,揭示了潜在空间中不一致推理的存在及早期规划阶段的安全监控重要性。

Comments 15 pages with 2 figures

Journal ref International Conference on Learning Representations (ICLR) Latent & Implicit Thinking (LIT) Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27484 2026-04-14 cs.LG cs.AI cs.CL 88%

Thought Branches: Interpreting LLM Reasoning Requires Resampling

思维分支:解释大语言模型推理需要重采样

Uzay Macar, Paul C. Bogdan, Senthooran Rajamanoharan, Neel Nanda

机构 * MATS

专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)

AI总结 本文通过重采样研究大语言模型推理过程中的因果影响,探讨了单个推理链的局限性,并提出了基于重采样的方法来分析模型决策和推理步骤的影响。

Comments Uzay Macar and Paul C. Bogdan contributed equally to this work, and their listed order was determined by coinflip

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06427 2026-04-09 cs.LG cs.AI cs.CL 88%

The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning

深度上限:大语言模型在发现潜在规划中的局限性

Yi Xu, Philipp Jettkant, Laura Ruis

机构 * University of Cambridge(剑桥大学) Imperial College London(帝国理工学院) MIT(麻省理工学院)

专题命中 规划推理 :planning(title,abstract);reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 研究探讨了大语言模型在无监督情况下发现多步规划策略的极限,发现模型在单次前向传递中能执行最多七步潜在规划,揭示了训练与测试时能力的差异。

Comments 10 pages, 3 figures, 1 table (30 pages, 9 figures, 10 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05112 2025-12-05 cs.CV cs.AI cs.CL cs.LG 88%

DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

DraCo: 文本到图像预览与稀有概念生成的草稿作为CoT

Dongzhi Jiang, Renrui Zhang, Haodong Li, Zhuofan Zong, Ziyu Guo, Jun He, Claire Guo, Junyan Ye, Rongyao Fang, Weijia Li, Rui Liu, Hongsheng Li

机构 * CUHK MMLab(香港中文大学 MMLab) CUHK IMIXR(香港中文大学 IMIXR) Sun Yat-Sen University(中山大学) SCUT(华南理工大学) CUHK (Shenzhen)(香港中文大学(深圳))

专题命中 规划推理 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);planning(abstract)

AI总结 DraCo通过结合文本和视觉内容的交错推理,提升文本到图像生成的精度和稀有概念生成能力。

Comments Project Page: https://github.com/CaraJ7/DraCo

详情

展开后加载摘要…

URL PDF HTML 收藏