arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-08-26 至 2025-08-26 共收录 120 信号源:cs.CL, cs.AI, cs.LG

1. 复杂问题求解 8 篇

2508.16729 2025-08-26 cs.CL 77%

Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?

Jason Li, Lauren Yraola, Kevin Zhu, Sean O'Brien

机构 * Algoverse AI Research(Algoverse AI 研究所)

专题命中 复杂问题求解 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Accepted to Insights @ NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23060 2025-08-26 cs.CL 70%

Self-Correcting Code Generation Using Small Language Models

Jeonghun Cho, Deokhyung Kang, Hyounghun Kim, Gary Geunbae Lee

机构 * Graduate School of Artificial Intelligence, POSTECH(人工智能研究生院,POSTECH) Department of Computer Science and Engineering, POSTECH(计算机科学与工程系,POSTECH)

专题命中 复杂问题求解 :reasoning(abstract);self-correction(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 (Findings, long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08925 2025-08-26 cs.LG cs.AI cs.CL 67%

Disentangling Exploration of Large Language Models by Optimal Exploitation

Tim Grams, Patrick Betz, Sascha Marton, Stefan Lüdtke, Christian Bartelt

机构 * Technical University of Clausthal(Clausthal 技术大学) University of Mannheim(曼海姆大学) University of Rostock(罗斯托克大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17330 2025-08-26 cs.CL cs.AI 62%

Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering

Boyuan Liu, Feng Ji, Jiayan Nan, Han Zhao, Weiling Chen, Shihao Xu, Xing Zhou

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15437 2025-08-26 cs.IR cs.AI cs.LG 62%

Test-time Corpus Feedback: From Retrieval to RAG

Mandeep Rathee, V Venktesh, Sean MacAvaney, Avishek Anand

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17565 2025-08-26 cs.AI 57%

TradingGroup: A Multi-Agent Trading System with Self-Reflection and Data-Synthesis

Feng Tian, Flora D. Salim, Hao Xue

机构 * The University of New South Wales(新南威尔士大学) University of New South Wales(新南威尔士大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12687 2025-08-26 cs.AI cs.CV 57%

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Virginia(弗吉尼亚大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06791 2025-08-26 cs.RO cs.AI cs.HC cs.MA 57%

AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social Robot

Xiao Wang, Lu Dong, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venugopal Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校) Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham(计算机科学与工程系,阿米特拉学校 computing,阿米特拉世界大学)

专题命中 复杂问题求解 :reasoning(abstract);分类 cs.AI

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 推理评测 35 篇

2508.16910 2025-08-26 cs.CL 85%

Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment

Bo Zhao, Yinghao Zhang, Ziqi Xu, Yongli Ren, Xiuzhen Zhang, Renqiang Luo, Zaiwen Feng, Feng Xia

机构 * Huazhong Agricultural University(华中农业大学) RMIT University(皇家墨尔本理工大学) Jilin University(吉林大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments This paper has been accepted to the 34th ACM International Conference on Information and Knowledge Management (CIKM 2025), Full Research Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16821 2025-08-26 cs.AI cs.LG 84%

PuzzleJAX: A Benchmark for Reasoning and Learning

Sam Earle, Graham Todd, Yuchen Li, Ahmed Khalifa, Muhammad Umair Nasir, Zehua Jiang, Andrzej Banburski-Fahey, Julian Togelius

机构 * New York University(纽约大学) University of Malta(马耳他大学) University of the Witwatersrand(沃斯兰大学) Microsoft(微软公司)

专题命中 推理评测 :reasoning(title,abstract);planning(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 11 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16998 2025-08-26 cs.CL cs.IR 83%

DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani, Adam Jatowt

机构 * University of Innsbruck(因斯布鲁克大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

Comments Accept at EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19599 2025-08-26 cs.SE cs.AI 83%

HoarePrompt: Structural Reasoning About Program Correctness in Natural Language

Dimitrios Stamatios Bouras, Yihan Dai, Tairan Wang, Yingfei Xiong, Sergey Mechtaev

机构 * Peking University(北京大学) Nankai University(南开大学) University College London(伦敦大学学院)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23266 2025-08-26 cs.CV cs.AI cs.CL 81%

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Ziyao Shangguan, Chuhan Li, Yuxuan Ding, Yanan Zheng, Yilun Zhao, Tesca Fitzgerald, Arman Cohan

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17258 2025-08-26 cs.CL cs.IR 79%

Are You Sure You're Positive? Consolidating Chain-of-Thought Agents with Uncertainty Quantification for Aspect-Category Sentiment Analysis

Filippos Ventirozos, Peter Appleby, Matthew Shardlow

机构 * Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) Autotrader Research Group(Autotrader 研究组) Autotrader UK(Autotrader 英国)

专题命中 推理评测 :chain-of-thought(title,abstract);分类 cs.CL

Comments 18 pages, 10 figures, 3 tables, Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025)

Journal ref Ventirozos et al. 2025. In Proc. of REALM 2025, pp. 309-326. ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13237 2025-08-26 eess.AS cs.CL cs.SD 79%

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Chih-Kai Yang, Neo Ho, Yen-Ting Piao, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

Comments Accepted to Interspeech 2025 (Oral). Update acknowledgement in this version. Project page: https://github.com/ckyang1124/SAKURA

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16861 2025-08-26 cs.CL 79%

Learning from Diverse Reasoning Paths with Routing and Collaboration

Zhenyu Lei, Zhen Tan, Song Wang, Yaochen Zhu, Zihan Chen, Yushun Dong, Jundong Li

机构 * University of Virginia(弗吉尼亚大学) Arizona State University(亚利桑那州立大学) Florida State University(佛罗里达州立大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16850 2025-08-26 cs.AI 79%

RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis

Anku Rani, Aparna Garimella, Apoorv Saxena, Balaji Vasan Srinivasan, Paul Pu Liang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16859 2025-08-26 cs.CV 78%

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li, Peipei Song, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (IAI), Hefei Comprehensive National Science Center(人工智能研究院(IAI),合肥综合性国家科学中心)

专题命中 推理评测 :reasoning(title,abstract)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17590 2025-08-26 cs.DB cs.AI cs.CL cs.MA 73%

RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System

Zui Chen, Han Li, Xinhao Zhang, Xiaoyu Chen, Chunyin Dong, Yifeng Wang, Xin Cai, Su Zhang, Ziqi Li, Chi Ding, Jinxu Li, Shuai Wang, Dousheng Zhao, Sanhai Gao, Guangyi Liu

机构 * Huawei Company(华为公司) Cornell University(康奈尔大学)

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 3 figures, 3 tables, to be submitted to VLDB 2026 (PVLDB Volume 19)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21805 2025-08-26 cs.CL cs.AI 73%

ImF: Implicit Fingerprint for Large Language Models

Jiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue, Juan Wen

机构 * China Agricultural University(中国农业大学)

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16987 2025-08-26 cs.AI cs.CV 70%

WebSight: A Vision-First Architecture for Robust Web Agents

Tanvir Bhathal, Asanshay Gupta

机构 * Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17580 2025-08-26 cs.CL cs.AI cs.LG 67%

UQ: Assessing Language Models on Unsolved Questions

Fan Nie, Ken Ziyu Liu, Zihao Wang, Rui Sun, Wei Liu, Weijia Shi, Huaxiu Yao, Linjun Zhang, Andrew Y. Ng, James Zou, Sanmi Koyejo, Yejin Choi, Percy Liang, Niklas Muennighoff

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments FN, KZL, and NM are project co-leads and contributed equally. Project website: https://uq.stanford.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17571 2025-08-26 cs.IR 67%

A Universal Framework for Offline Serendipity Evaluation in Recommender Systems via Large Language Models

Yu Tokutake, Kazushi Okamoto, Kei Harada, Atsushi Shibata, Koki Karube

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract)

Journal ref The 34th ACM International Conference on Information and Knowledge Management (CIKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17290 2025-08-26 cs.AI cs.LG 62%

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment

Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah

机构 * Computer Engineering Department, Sharif University of Technology, Iran(谢尔盖大学计算机工程系,伊朗) Qatar Computing Research Institute, Qatar(卡塔尔计算研究所,卡塔尔) Computer Engineering Department, University of Isfahan, Iran(伊斯法罕大学计算机工程系,伊朗) Independent Researcher(独立研究者)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11280 2025-08-26 cs.CL cs.AI 62%

LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought

Ruiyan Qi, Congding Wen, Weibo Zhou, Jiwei Li, Shangsong Liang, Lingbo Li

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01881 2025-08-26 cs.AI cs.CL 62%

WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue

Yaoyao Qian, Jindan Huang, Yuanli Wang, Simon Yu, Kyrie Zhixuan Zhou, Jiayuan Mao, Mingfu Liang, Hanhan Zhou

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Project website: https://nanostorm.netlify.app/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15587 2025-08-26 cs.CL cs.AI cs.IR 62%

SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain

Dakuan Lu, Xiaoyu Tan, Rui Xu, Tianchu Yao, Chao Qu, Wei Chu, Yinghui Xu, Yuan Qi

机构 * INFLY TECH (Shanghai) Co., Ltd.(INFLY科技(上海)有限公司) Fudan University(复旦大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18179 2025-08-26 cs.AI cs.CV 57%

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Zhenwei Tang, Difan Jiao, Blair Yang, Ashton Anderson

机构 * Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Coolwei AI Lab(Coolwei人工智能实验室)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18040 2025-08-26 cs.AI 57%

PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration

Xin Wang, Zhiyao Cui, Hao Li, Ya Zeng, Chenxu Wang, Ruiqi Song, Yihang Chen, Kun Shao, Qiaosheng Zhang, Jinzhuo Liu, Siyue Ren, Shuyue Hu, Zhen Wang

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17703 2025-08-26 cs.CL 57%

EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning

Yinda Chen, Yangfan He, Jing Yang, Dapeng Zhang, Zhenlong Yuan, Muhammad Attique Khan, Jamel Baili, Por Lip Yee

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Department of Computer Science, University of Minnesota-Twin Cities(计算机科学系,明尼苏达大学双城分校) Center of Research for Cyber Security and Network (CSNET), Faculty of Computer Science and Information Technology, Universiti Malaya(网络安全与网络研究中心(CSNET),马来亚大学计算机科学与信息技术学院) DSLAB, School of Information Science & Engineering, Lanzhou University(信息科学与工程学院,兰州大学) Institute of Computing Technology, Chinese Academy of Sciences(计算技术研究所,中国科学院) Department of AI, Prince Mohammad bin Fahd University(人工智能系,普林姆·法赫德大学) Department of Computer Engineering, College of Computer Science, King Khalid University(计算机工程系,国王·卡利德大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏