arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-09-18 至 2025-09-18 共收录 50 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 13 篇

2509.01081 2025-09-18 cs.CL cs.AI 81%

Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation

Abdessalam Bouchekif, Samer Rashwani, Heba Sbahi, Shahd Gaben, Mutaz Al-Khatib, Mohammed Ghaly

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 7 Tables, Code: https://github.com/bouchekif/inheritance_evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19676 2025-09-18 cs.AI 80%

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

Lachlan McGinness, Peter Baumgartner

机构 * School of Computer Science, Australian National University and CSIRO(计算机科学学院,澳大利亚国立大学和CSIRO)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments The original version of this article was withdrawn because there were errors in the evaluation of model faithfulness to reasoning strategies and completeness of reasoning. The analysis was re-conducted correctly and version two contains the corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08364 2025-09-18 cs.AI 79%

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

Enci Zhang, Xingang Yan, Wei Lin, Tianxiang Zhang, Qianchun Lu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments 14 pages, 3 figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13941 2025-09-18 cs.SE cs.AI cs.CL 62%

An Empirical Study on Failures in Automated Issue Solving

Simiao Liu, Fang Liu, Liehao Li, Xin Tan, Yinghao Zhu, Xiaoli Lian, Li Zhang

机构 * Beihang University(北京航空航天大学) The University of Hong Kong(香港大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13773 2025-09-18 cs.AI cs.IR 57%

MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation

Zhipeng Bian, Jieming Zhu, Xuyang Xie, Quanyu Dai, Zhou Zhao, Zhenhua Dong

机构 * Shenzhen University(深圳大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21137 2025-09-18 cs.CL 57%

How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations

Yoshiki Takenami, Yin Jou Huang, Yugo Murawaki, Chenhui Chu

机构 * Kyoto University(京都大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 18 pages, 2 figures. Accepted to EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 57%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05262 2025-09-18 cs.CL 57%

Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic

Yang Yan, Yu Lu, Renjun Xu, Zhenzhong Lan

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14227 2025-09-18 cs.CV 50%

Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark

Nisarg A. Shah, Amir Ziai, Chaitanya Ekanadham, Vishal M. Patel

机构 * Netflix, Inc.(Netflix公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理评测 :reasoning(abstract)

Comments 11 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13691 2025-09-18 cs.RO 50%

SPAR: Scalable LLM-based PDDL Domain Generation for Aerial Robotics

Songhao Huang, Yuwei Wu, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar

机构 * GRASP Lab, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 推理评测 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13572 2025-09-18 cs.RO 50%

Using Visual Language Models to Control Bionic Hands: Assessment of Object Perception and Grasp Inference

Ozan Karaali, Hossam Farag, Strahinja Dosen, Cedomir Stefanovic

机构 * Department of Electronic Systems, Aalborg University, Denmark(电子系统系,奥胡斯大学) Department of Health Science and Technology, Aalborg University, Denmark(健康科学与技术系,奥胡斯大学)

专题命中 推理评测 :planning(abstract)

Comments ICAT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他推理 9 篇

2509.13683 2025-09-18 cs.CL cs.AI 81%

Improving Context Fidelity via Native Retrieval-Augmented Reasoning

Suyuchen Wang, Jinlin Wang, Xinyu Wang, Shiqi Li, Xiangru Tang, Sirui Hong, Xiao-Wen Chang, Chenglin Wu, Bang Liu

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10408 2025-09-18 cs.LG cs.CL 81%

Out-of-Context Reasoning in Large Language Models

Jonathan Shaki, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus

机构 * Bar-Ilan University(巴伊兰大学) University of Oxford(牛津大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08627 2025-09-18 cs.SE 67%

NL in the Middle: Code Translation with LLMs and Intermediate Representations

Chi-en Amy Tai, Pengyu Nie, Lukasz Golab, Alexander Wong

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19546 2025-09-18 cs.CL cs.AI 62%

Language Models Identify Ambiguities and Exploit Loopholes

Jio Choi, Mohit Bansal, Elias Stengel-Eskin

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 camera-ready; Code: https://github.com/esteng/ambiguous-loophole-exploitation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14034 2025-09-18 cs.CL 57%

Enhancing Multi-Agent Debate System Performance via Confidence Expression

Zijie Lin, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.CL

Comments EMNLP'25 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19505 2025-09-18 cs.AI 57%

Caught in the Act: a mechanistic approach to detecting deception

Gerard Boxo, Ryan Socha, Daniel Yoo, Shivam Raval

机构 * Barcelona Institute of Science and Technology(巴塞罗那科学与技术研究所) NorthWest Arkansas Community College(西北阿肯色社区学院) Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学)

专题命中 其他推理 :reasoning(abstract);分类 cs.AI

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14187 2025-09-18 eess.AS 50%

Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs

Yu-Wen Chen, Melody Ma, Julia Hirschberg

专题命中 其他推理 :reasoning(abstract)

Comments EMNLP 2025 MainConference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10676 2025-09-18 hep-ex 50%

NuGraph2 with Explainability: Post-hoc Explanations for Geometric Neural Network Predictions

Margaret Voetberg, Vitor F. Grizzi, Giuseppe Cerati, Hadi Meidani, V Hewes

专题命中 其他推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13919 2025-09-18 cs.CV 50%

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

Yuanchen Wu, Ke Yan, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li

机构 * School of Computer Engineering(计算机工程学院) Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) Efficient Computing, Xiamen University(高效计算,厦门大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 其他推理 :reasoning(abstract)

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏