arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-09-18 至 2025-09-18 共收录 13 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 13 篇

2509.14180 2025-09-18 cs.CL cs.AI cs.LG 85%

Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs

Akhil Theerthala

机构 * Perfios Software Solutions(Perfios软件解决方案)

专题命中 推理评测 :reasoning(title,abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 24 pages, 11 figures. The paper presents a novel framework for generating a personal finance dataset. The resulting fine-tuned model and dataset are publicly available

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05170 2025-09-18 cs.SE cs.AI cs.CL cs.LG 82%

Posterior-GRPO: Rewarding Reasoning Processes in Code Generation

Lishui Fan, Yu Zhang, Mouxiang Chen, Zhongxin Liu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01081 2025-09-18 cs.CL cs.AI 81%

Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation

Abdessalam Bouchekif, Samer Rashwani, Heba Sbahi, Shahd Gaben, Mutaz Al-Khatib, Mohammed Ghaly

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 7 Tables, Code: https://github.com/bouchekif/inheritance_evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19676 2025-09-18 cs.AI 80%

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

Lachlan McGinness, Peter Baumgartner

机构 * School of Computer Science, Australian National University and CSIRO(计算机科学学院,澳大利亚国立大学和CSIRO)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments The original version of this article was withdrawn because there were errors in the evaluation of model faithfulness to reasoning strategies and completeness of reasoning. The analysis was re-conducted correctly and version two contains the corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08364 2025-09-18 cs.AI 79%

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

Enci Zhang, Xingang Yan, Wei Lin, Tianxiang Zhang, Qianchun Lu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

Comments 14 pages, 3 figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13941 2025-09-18 cs.SE cs.AI cs.CL 62%

An Empirical Study on Failures in Automated Issue Solving

Simiao Liu, Fang Liu, Liehao Li, Xin Tan, Yinghao Zhu, Xiaoli Lian, Li Zhang

机构 * Beihang University(北京航空航天大学) The University of Hong Kong(香港大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13773 2025-09-18 cs.AI cs.IR 57%

MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation

Zhipeng Bian, Jieming Zhu, Xuyang Xie, Quanyu Dai, Zhou Zhao, Zhenhua Dong

机构 * Shenzhen University(深圳大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), ACL 2025. Official version: https://doi.org/10.18653/v1/2025.acl-industry.103

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) ACL 2025 1457-1465

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21137 2025-09-18 cs.CL 57%

How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations

Yoshiki Takenami, Yin Jou Huang, Yugo Murawaki, Chenhui Chu

机构 * Kyoto University(京都大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments 18 pages, 2 figures. Accepted to EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 57%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05262 2025-09-18 cs.CL 57%

Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic

Yang Yan, Yu Lu, Renjun Xu, Zhenzhong Lan

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

Comments Accepted by EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14227 2025-09-18 cs.CV 50%

Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark

Nisarg A. Shah, Amir Ziai, Chaitanya Ekanadham, Vishal M. Patel

机构 * Netflix, Inc.(Netflix公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理评测 :reasoning(abstract)

Comments 11 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13691 2025-09-18 cs.RO 50%

SPAR: Scalable LLM-based PDDL Domain Generation for Aerial Robotics

Songhao Huang, Yuwei Wu, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar

机构 * GRASP Lab, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 推理评测 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13572 2025-09-18 cs.RO 50%

Using Visual Language Models to Control Bionic Hands: Assessment of Object Perception and Grasp Inference

Ozan Karaali, Hossam Farag, Strahinja Dosen, Cedomir Stefanovic

机构 * Department of Electronic Systems, Aalborg University, Denmark(电子系统系,奥胡斯大学) Department of Health Science and Technology, Aalborg University, Denmark(健康科学与技术系,奥胡斯大学)

专题命中 推理评测 :planning(abstract)

Comments ICAT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏