arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45417 信号源:cs.CL, cs.AI, cs.LG

1. 数学推理 3252 篇

2304.09102 2023-04-19 cs.CL cs.AI 62%

Solving Math Word Problems by Combining Language Models With Symbolic Solvers

Joy He-Yueya, Gabriel Poesia, Rose E. Wang, Noah D. Goodman

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11436 2023-04-13 cs.CL cs.AI 62%

Mind meets machine: Unravelling GPT-4's cognitive psychology

Sifatkaur Dhingra, Manmeet Singh, Vaisakh SB, Neetiraj Malviya, Sukhpal Singh Gill

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02015 2023-04-06 cs.CL cs.AI 62%

How well do Large Language Models perform in Arithmetic tasks?

Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang

专题命中 数学推理 :chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07382 2023-02-14 cs.CL cs.AI 62%

Behavior Cloned Transformers are Neurosymbolic Reasoners

Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, Prithviraj Ammanabrolu

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12835 2022-11-24 cs.CL cs.CY cs.LG 62%

Automatic Generation of Socratic Subquestions for Teaching Math Word Problems

Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, Mrinmaya Sachan

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.LG

Comments Kumar Shridhar and Jakub Macina contributed equally to this work. Accepted at the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022). Code available: https://github.com/eth-nlped/scaffolding-generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11522 2022-10-24 cs.CV cs.AI cs.LG 62%

Composing Ensembles of Pre-trained Models via Iterative Consensus

Shuang Li, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, Igor Mordatch

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.15594 2022-09-14 cs.LG cs.AI 62%

A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level

Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, Roman Wang, Nikhil Singh, Taylor L. Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, Gilbert Strang

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 181 pages, 8 figures, 280 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12255 2022-05-25 cs.CL cs.AI 62%

TALM: Tool Augmented Language Models

Aaron Parisi, Yao Zhao, Noah Fiedel

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10382 2022-04-25 cs.AI cs.LG 62%

Facilitating automated conversion of scientific knowledge into scientific simulation models with the Machine Assisted Generation, Calibration, and Comparison (MAGCC) Framework

Chase Cockrell, Scott Christley, Gary An

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 20 Pages, 8 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02942 2022-02-08 cs.AI cs.CC cs.LG cs.LO 62%

Tractable Boolean and Arithmetic Circuits

Adnan Darwiche

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments An earlier version of this article appeared in the following edited book. Pascal Hitzler and Md Kamruzzaman Sarker, editors. Neuro-Symbolic Artificial Intelligence: The State of the Art, volume 342 of Frontiers in Artificial Intelligence and Applications. IOS Press, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11446 2022-01-24 cs.CL cs.AI 62%

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d'Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, Geoffrey Irving

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 120 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04255 2021-12-17 cs.AI cs.CL cs.IR 62%

Quantum Mathematics in Artificial Intelligence

Dominic Widdows, Kirsty Kitto, Trevor Cohen

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Adding journal reference, recommended by JAIR editors upon publication

Journal ref Journal of Artificial Intelligence Research 72 (2021) 1307-1341

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03137 2021-10-14 cs.CL cs.LG 62%

NumGPT: Improving Numeracy Ability of Generative Pre-trained Models

Zhihua Jin, Xin Jiang, Xingbo Wang, Qun Liu, Yong Wang, Xiaozhe Ren, Huamin Qu

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.LG

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06196 2021-09-10 cs.CL cs.AI 62%

Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations

Tianqiao Liu, Qiang Fang, Wenbiao Ding, Hang Li, Zhongqin Wu, Zitao Liu

专题命中 数学推理 :planning(abstract);分类 cs.CL、cs.AI

Comments EMNLP'21: The 2021 Conference on Empirical Methods in Natural Language Processing, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14953 2021-03-30 cs.CL cs.AI cs.IT math.IT 62%

What they do when in doubt: a study of inductive biases in seq2seq learners

Eugene Kharitonov, Rahma Chaabouni

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.06954 2020-01-14 cs.AI cs.LG cs.LO 62%

Semiring Programming: A Declarative Framework for Generalized Sum Product Problems

Vaishak Belle, Luc De Raedt

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments In AAAI Workshop: Statistical Relational Artificial Intelligence, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.00509 2015-08-12 cs.AI cs.LG 62%

Maintaining prediction quality under the condition of a growing knowledge space

Christoph Jahnz

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11612 2025-11-18 cs.DC cs.AI 61%

Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems

Aasish Kumar Sharma, Julian Kunkel

机构 * Faculty of Mathematics and Computer Science, Georg-August-Universität Göttingen(数学与计算机科学学院,哥廷根乔治-奥古斯特大学)

专题命中 数学推理 :reasoning(abstract,comments);分类 cs.AI

Comments 14 pages, 4 figures, 2 tables. Evaluation study on LLM-based reasoning for HPC scheduling. Published in Research in Academic Engineering Journal (RAEJ), 2025

Journal ref Robot Autom Eng J. 2025; 6(5): 555696

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17249 2023-07-03 cs.NE cs.AI 61%

A Hybrid System for Systematic Generalization in Simple Arithmetic Problems

Flavio Petruzzellis, Alberto Testolin, Alessandro Sperduti

专题命中 数学推理 :reasoning(abstract,comments);分类 cs.AI

Comments Accepted at NeSy 2023, 17th International Workshop on Neural-Symbolic Learning and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25992 2026-08-27 cs.AI cs.MA 新提交 57%

ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs

ProgRouter:面向质量-成本权衡的多智能体大语言模型工作流的在线进度引导编排

Somgyuan Li, Ahmed M. Abdelmoniem, Shiqiang Wang

机构 * Aston University(阿斯顿大学) Queen Mary University of London(伦敦玛丽女王大学) University of Exeter(埃克塞特大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 ProgRouter是一种在线进度引导路由框架,通过多视图任务进度评分器等机制,在多智能体LLM工作流中平衡任务质量与时间、成本预算,在多类任务数据集上较基线降低运营成本且保持性能。

Comments Accepted in Findings of the Association for Computational Linguistics: EMNLP 2026. Index Terms: Collaborative agentic workflows, LLM agent orchestration, Quality-cost trade-off, Task progress prediction, Online decision-making

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25311 2026-08-27 cs.LG 新提交 57%

Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models

前缀去噪一致性:扩散语言模型的测试时验证方法

Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) Nagoya University(名古屋大学) RIKEN AIP(理化学研究所人工智能项目)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 该研究针对扩散语言模型提出测试时自验证方法PDC,通过前缀条件再生的轨迹稳定性差异修正初始生成样本,在两类推理基准中提升性能且具鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23867 2026-08-26 cs.MA cs.CL 新提交 57%

Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information

市场而非规划者:利用私有信息对大语言模型智能体进行去中心化协调

Xiao Liu, Haoyang Li, Songwei Li, Hongbo Fang, Fengli Xu, Feng Shi, James Evans

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 针对中心化LLM智能体协调的瓶颈与易操控问题,提出去中心化的AgentLance重复劳动力市场机制,经多类任务验证,其匹配专长、转向廉价智能体的表现优于基线方法,还可通过修正市场失灵进一步提升效率。

Comments Working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23632 2026-08-26 cs.AI cs.SE 新提交 57%

Function-Level Execution Feedback for Code Preference Optimization

用于代码偏好优化的函数级执行反馈

Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee

机构 * Seoul National University(首尔大学) Konkuk University(建国大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出STEP-KTODER框架,将代码步骤定义为多函数程序的模块级函数,结合函数级过程监督与程序结果级反馈,在多个代码基准测试中优于KTO、DPO等方法,且验证了执行标签的重要性。

Comments 20 pages, 8 figures, 14 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23311 2026-08-26 cs.CL 版本更新 57%

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

超越稳定性-探索困境:面向大语言模型策略优化的环境正则化方法

Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu

机构 * Alibaba Group(阿里巴巴集团) Beijing Normal University(北京师范大学) National University of Singapore(新加坡国立大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 该研究针对大语言模型策略优化的稳定性-探索困境,提出ERPO方法,通过查询KL散度正则化输入侧分布,接入现有策略优化流水线,在数学推理基准上实现了更强准确率与更稳定行为。

Comments Accepted to EMNLP 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14673 2026-08-26 cs.AI cs.GL quant-ph 版本更新 57%

Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument

AI生成数学证明的审计:量子并行重复中一个贪心条件引理的修正

Mikołaj Sienicki, Krzysztof Sienicki

机构 * Polish–Japanese Academy of Information Technology(波兰-日本信息技术学院) Chair of Theoretical Physics of Naturally Intelligent Systems (NIS)(自然智能系统理论物理研究所(NIS))

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本文审计OpenAI相关工作中AI生成的数学证明,发现其量子并行重复中贪心条件引理的证明存在极性错误,给出反例并提供局部修正证明,揭示AI论证可能隐藏互补事件的关键反转。

Comments 8 pages, 5 references. V2: 11 pages, 8 references. Substantially revised. The polarity error reported in v1 is withdrawn: it resulted from loss of an overbar during PDF text extraction, not from an error in the OpenAI manuscript. The revised paper retains and extends the independent human assessment of Chapter 6. No confirmed mathematical error was found in the examined argument

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22806 2026-08-25 cs.CL 新提交 57%

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

DIAG:面向数据高效数学偏好蒸馏的诊断式迭代对齐与生成

Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu

机构 * Tsinghua University(清华大学) Kyoto University(京都大学) Nanjing University(南京大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 DIAG是一种自适应调整训练分布的框架,通过诊断偏好对产出与生成针对性训练数据,提升数学推理任务中偏好监督的信息价值,在同等训练预算下增强模型性能。

Comments Accepted by EMNLP 2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22646 2026-08-25 cs.AI 新提交 57%

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

CAI-DLLM:面向扩散语言模型的感知收敛推理

Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 CAI-DLLM 是一种无需训练的扩散语言模型推理方法,通过利用第一步置信度调整解码策略,在多类任务上实现显著推理加速,同时保持较高准确率并降低能耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21496 2026-08-25 cs.LG math.ST stat.ML stat.TH 新提交 57%

The geometry of AI validation: Exact certification limits for iid best-of-N search

AI验证的几何:独立同分布N选最优搜索的精确认证极限

Ricardo Fitas

机构 * Technical University of Darmstadt(达姆施塔特工业大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 该研究针对独立同分布N选最优搜索推导了精确认证极限,提出双门审计规则,其在数学推理和代码选择的回顾性分析中可降低保留误差。

Comments 32 pages, 4 figures; includes Supporting Information. Code and processed data are available at https://github.com/rfitas-lab/geometry-of-ai-validation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14685 2026-08-25 cs.LG stat.ML 版本更新 57%

Rethinking Reverse KL as Adaptive Entropy Distillation

将反向KL重新思考为自适应熵蒸馏

Shizhen Li, Zhiyu Shen, Yuyin Lu, Yunhe Pang, Jielin Song, Yanghui Rao, Fu Lee Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 该研究针对知识蒸馏难以平衡忠实模仿与鲁棒生成的问题,提出自适应熵蒸馏(AED)方法,通过分解反向KL目标函数并利用教师熵动态校准模仿强度,在指令遵循与数学推理基准上表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00869 2026-08-25 cs.LG 版本更新 57%

Enhancing LLM Metacognition via Cognitive Pairwise Training

通过认知成对训练增强LLM元认知

Weitao Li, Hao Zhou, Xuanyu Lei, Fandong Meng, Yuanhang Liu, Jingyi Ren, Ante Wang, Xiaolong Wang, Yuanchi Zhang, Fuwen Luo, Guangwen Yang, Lin Gan, Weizhi Ma, Yang Liu

机构 * National Engineering Laboratory for Intelligent Information Processing, Academy of Mathematics and Physics, Chinese Academy of Sciences(智能信息处理国家工程实验室,中国科学院数学物理研究所) University of Science and Technology of China(中国科学技术大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 提出认知成对训练(CPT),通过成对比较推理轨迹来学习区分可靠与不可靠推理,从而提升LLM的推理与元认知权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏