arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-10-24 至 2025-10-24 共收录 81 信号源:cs.CL, cs.AI, cs.LG

1. 规划推理 17 篇

2505.23010 2025-10-24 cs.CV 50%

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model

Bowen Chen, Keyan Chen, Mohan Yang, Zhengxia Zou, Zhenwei Shi

专题命中 规划推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19868 2025-10-24 cs.SE 50%

Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation

Qian Xiong, Bo Yang, Weisong Sun, Yiran Zhang, Tianlin Li, Yang Liu, Zhi Jin

专题命中 规划推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉空间推理 6 篇

2506.05341 2025-10-24 cs.CV cs.AI 87%

Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning

Xingjian Ran, Yixuan Li, Linning Xu, Mulin Yu, Bo Dai

机构 * The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉空间推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)

Comments Project Page: https://directlayout.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20696 2025-10-24 cs.CV 85%

Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward

Jing Bi, Guangyu Sun, Ali Vosoughi, Chen Chen, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Central Florida(中央佛罗里达大学)

专题命中 视觉空间推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20198 2025-10-24 cs.CL cs.AI 81%

Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models

Maggie Bai, Ava Kim Cohen, Eleanor Koss, Charlie Lichtenbaum

专题命中 视觉空间推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 20 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20161 2025-10-24 cs.RO 67%

PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation

Ahmed Alanazi, Duy Ho, Yugyung Lee

机构 * Department of Computer Science, University of Missouri–Kansas City (UMKC)(密苏里大学哥伦比亚分校计算机科学系) Department of Computer Science, California State University, Fullerton(加州州立大学富尔顿分校计算机科学系)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract)

Comments 8 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 57%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20385 2025-10-24 cs.CV 50%

Positional Encoding Field

Yunpeng Bai, Haoxiang Li, Qixing Huang

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Pixocial Technology(Pixocial技术)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 测试时计算 7 篇

2505.15034 2025-10-24 cs.LG cs.AI cs.CL 89%

RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning

Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong, Duane S. Boning, Dina Katabi

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

专题命中 测试时计算 :reasoning(title,abstract);verifier(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18454 2025-10-24 cs.CL 85%

Hybrid Latent Reasoning via Reinforcement Learning

Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, Dong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Google(谷歌) LMU(莱比锡马克斯·普朗克研究所)

专题命中 测试时计算 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04210 2025-10-24 cs.AI cs.CL 81%

Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models

Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu, Mengdi Wang, Dinesh Manocha, Furong Huang, Mohammad Ghavamzadeh, Amrit Singh Bedi

机构 * University of Maryland(马里兰大学) Princeton University(普林斯顿大学) Capital One(Capital One公司) Amazon AGI University of Central Florida(佛罗里达中央大学)

专题命中 测试时计算 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15275 2025-10-24 cs.AI cs.LG 81%

Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning

Jie Cheng, Gang Xiong, Ruixi Qiao, Lijun Li, Chao Guo, Junle Wang, Yisheng Lv, Fei-Yue Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tencent(腾讯) Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院) DeSci Center of Parallel Intelligence, Obuda University(并行智能德Sci中心,奥布达大学) Research Center for Chinese Economics and Social Security, University of Chinese Academy of Sciences(中国经济社会安全研究中心,中国科学院大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 测试时计算 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17472 2025-10-24 stat.ML cs.LG 79%

Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs

Paula Cordero-Encinar, Andrew B. Duncan

专题命中 测试时计算 :reasoning(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 62%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14827 2025-10-24 cs.CL cs.AI 62%

Text Generation Beyond Discrete Token Sampling

Yufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang, Jianfeng Gao

机构 * UC San Diego(加州大学圣地亚哥分校) Microsoft Research(微软研究院)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 复杂问题求解 13 篇

2505.14667 2025-10-24 cs.AI cs.CL 87%

SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment

Wonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert No

机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) Department of Computer Science and Engineering, Yonsei University(计算机科学与工程系,延世大学)

专题命中 复杂问题求解 :reasoning(title,abstract);chain-of-thought(title);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025. Code and models are available at https://ai-isl.github.io/safepath

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22651 2025-10-24 cs.CV cs.CL cs.LG 86%

Sherlock: Self-Correcting Reasoning in Vision-Language Models

Yi Ding, Ruqi Zhang

机构 * Department of Computer Science, Purdue University, USA(计算机科学系,普渡大学)

专题命中 复杂问题求解 :reasoning(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.LG

Comments Published at NeurIPS 2025, 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18254 2025-10-24 cs.AI cs.LG 84%

Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning

Sion Weatherhead, Flora Salim, Aaron Belbasis

机构 * University of New South Wales Sydney(新南威尔士大学悉尼分校) Aurecon Group(Aurecon集团)

专题命中 复杂问题求解 :reasoning(title,abstract);self-correction(abstract);分类 cs.AI、cs.LG

Comments Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20780 2025-10-24 cs.CL cs.AI 81%

Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost

Runzhe Zhan, Zhihong Huang, Xinyi Yang, Lidia S. Chao, Min Yang, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20607 2025-10-24 cs.LG cs.AI 81%

Generalizable Reasoning through Compositional Energy Minimization

Alexandru Oarga, Yilun Du

机构 * University of Barcelona(巴塞罗那大学) Harvard University(哈佛大学)

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20188 2025-10-24 cs.AI 79%

TRUST: A Decentralized Framework for Auditing Large Language Model Reasoning

Morris Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li, Zhuo Zhang, Tianlong Chen

专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04462 2025-10-24 cs.CL cs.AI 73%

Benchmarking GPT-5 for biomedical natural language processing

Yu Hou, Zaifu Zhan, Min Zeng, Yifan Wu, Shuang Zhou, Rui Zhang

专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20653 2025-10-24 stat.ML cs.AI cs.LG 66%

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

Jack Butler, Nikita Kozodoi, Zainab Afolabi, Brian Tyacke, Gaiar Baimuratov

机构 * Amazon Web Services(亚马逊网络服务) Zalando

专题命中 复杂问题求解 :reasoning(abstract,journal_ref);分类 cs.AI、cs.LG

Journal ref Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20807 2025-10-24 cs.CV cs.LG 57%

Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers

Dean L Slack, G Thomas Hudson, Thomas Winterbottom, Noura Al Moubayed

机构 * Durham University(杜伦大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.LG

Comments 14 pages, 14 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, 36, 19106-19118, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20637 2025-10-24 cs.LG 57%

Large Multimodal Models-Empowered Task-Oriented Autonomous Communications: Design Methodology and Implementation Challenges

Hyun Jong Yang, Hyunsoo Kim, Hyeonho Noh, Seungnyun Kim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Hanbat National University(翰baum国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11851 2025-10-24 cs.CR cs.CL 57%

Deep Research Brings Deeper Harm

Shuo Chen, Zonggen Li, Zhen Han, Bailan He, Tong Liu, Haokun Chen, Georg Groh, Philip Torr, Volker Tresp, Jindong Gu

机构 * LMU Munich(慕尼黑大学) Siemens(西门子) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Technical University of Munich (TUM)(慕尼黑技术大学) AWS AI(亚马逊AI) Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse卓越可靠AI学校) University of Hong Kong (HKU)(香港大学) University of Oxford(牛津大学)

专题命中 复杂问题求解 :planning(abstract);分类 cs.CL

Comments Accepted to Reliable ML from Unreliable Data Workshop @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20622 2025-10-24 cs.CV 50%

SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding

Yuan Sheng, Yanbin Hao, Chenxu Li, Shuo Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

专题命中 复杂问题求解 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20223 2025-10-24 cs.CR cs.MM 50%

Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations

Divyanshu Kumar, Shreyas Jena, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 复杂问题求解 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 推理评测 17 篇

2510.20815 2025-10-24 cs.IR 82%

Generative Reasoning Recommendation via LLMs

Minjie Hong, Zetong Zhou, Zirun Guo, Ziang Zhang, Ruofan Hu, Weinan Gan, Jieming Zhu, Zhou Zhao

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09933 2025-10-24 cs.AI cs.CL cs.LG 82%

MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?

Kai Yan, Zhan Ling, Kang Liu, Yifan Yang, Ting-Han Fan, Lingfeng Shen, Zhengyin Du, Jiecao Chen

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 39 pages, 11 figures. The paper is accepted at NeurIPS 2025 Datasets & Benchmarks Track, and the latest version adds modifications in camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏