arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3246 信号源:cs.CL, cs.AI, cs.LG

1. 数学推理 3246 篇

2211.12835 2022-11-24 cs.CL cs.CY cs.LG 62%

Automatic Generation of Socratic Subquestions for Teaching Math Word Problems

Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, Mrinmaya Sachan

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.LG

Comments Kumar Shridhar and Jakub Macina contributed equally to this work. Accepted at the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022). Code available: https://github.com/eth-nlped/scaffolding-generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11522 2022-10-24 cs.CV cs.AI cs.LG 62%

Composing Ensembles of Pre-trained Models via Iterative Consensus

Shuang Li, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, Igor Mordatch

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.15594 2022-09-14 cs.LG cs.AI 62%

A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level

Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, Roman Wang, Nikhil Singh, Taylor L. Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, Gilbert Strang

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 181 pages, 8 figures, 280 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12255 2022-05-25 cs.CL cs.AI 62%

TALM: Tool Augmented Language Models

Aaron Parisi, Yao Zhao, Noah Fiedel

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10382 2022-04-25 cs.AI cs.LG 62%

Facilitating automated conversion of scientific knowledge into scientific simulation models with the Machine Assisted Generation, Calibration, and Comparison (MAGCC) Framework

Chase Cockrell, Scott Christley, Gary An

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 20 Pages, 8 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02942 2022-02-08 cs.AI cs.CC cs.LG cs.LO 62%

Tractable Boolean and Arithmetic Circuits

Adnan Darwiche

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments An earlier version of this article appeared in the following edited book. Pascal Hitzler and Md Kamruzzaman Sarker, editors. Neuro-Symbolic Artificial Intelligence: The State of the Art, volume 342 of Frontiers in Artificial Intelligence and Applications. IOS Press, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11446 2022-01-24 cs.CL cs.AI 62%

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d'Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, Geoffrey Irving

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 120 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04255 2021-12-17 cs.AI cs.CL cs.IR 62%

Quantum Mathematics in Artificial Intelligence

Dominic Widdows, Kirsty Kitto, Trevor Cohen

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Adding journal reference, recommended by JAIR editors upon publication

Journal ref Journal of Artificial Intelligence Research 72 (2021) 1307-1341

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03137 2021-10-14 cs.CL cs.LG 62%

NumGPT: Improving Numeracy Ability of Generative Pre-trained Models

Zhihua Jin, Xin Jiang, Xingbo Wang, Qun Liu, Yong Wang, Xiaozhe Ren, Huamin Qu

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.LG

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06196 2021-09-10 cs.CL cs.AI 62%

Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations

Tianqiao Liu, Qiang Fang, Wenbiao Ding, Hang Li, Zhongqin Wu, Zitao Liu

专题命中 数学推理 :planning(abstract);分类 cs.CL、cs.AI

Comments EMNLP'21: The 2021 Conference on Empirical Methods in Natural Language Processing, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14953 2021-03-30 cs.CL cs.AI cs.IT math.IT 62%

What they do when in doubt: a study of inductive biases in seq2seq learners

Eugene Kharitonov, Rahma Chaabouni

专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.06954 2020-01-14 cs.AI cs.LG cs.LO 62%

Semiring Programming: A Declarative Framework for Generalized Sum Product Problems

Vaishak Belle, Luc De Raedt

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments In AAAI Workshop: Statistical Relational Artificial Intelligence, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.00509 2015-08-12 cs.AI cs.LG 62%

Maintaining prediction quality under the condition of a growing knowledge space

Christoph Jahnz

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11612 2025-11-18 cs.DC cs.AI 61%

Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems

Aasish Kumar Sharma, Julian Kunkel

机构 * Faculty of Mathematics and Computer Science, Georg-August-Universität Göttingen(数学与计算机科学学院,哥廷根乔治-奥古斯特大学)

专题命中 数学推理 :reasoning(abstract,comments);分类 cs.AI

Comments 14 pages, 4 figures, 2 tables. Evaluation study on LLM-based reasoning for HPC scheduling. Published in Research in Academic Engineering Journal (RAEJ), 2025

Journal ref Robot Autom Eng J. 2025; 6(5): 555696

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17249 2023-07-03 cs.NE cs.AI 61%

A Hybrid System for Systematic Generalization in Simple Arithmetic Problems

Flavio Petruzzellis, Alberto Testolin, Alessandro Sperduti

专题命中 数学推理 :reasoning(abstract,comments);分类 cs.AI

Comments Accepted at NeSy 2023, 17th International Workshop on Neural-Symbolic Learning and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23867 2026-08-26 cs.MA cs.CL 新提交 57%

Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information

市场而非规划者:利用私有信息对大语言模型智能体进行去中心化协调

Xiao Liu, Haoyang Li, Songwei Li, Hongbo Fang, Fengli Xu, Feng Shi, James Evans

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 针对中心化LLM智能体协调的瓶颈与易操控问题,提出去中心化的AgentLance重复劳动力市场机制,经多类任务验证,其匹配专长、转向廉价智能体的表现优于基线方法,还可通过修正市场失灵进一步提升效率。

Comments Working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23632 2026-08-26 cs.AI cs.SE 新提交 57%

Function-Level Execution Feedback for Code Preference Optimization

用于代码偏好优化的函数级执行反馈

Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee

机构 * Seoul National University(首尔大学) Konkuk University(建国大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出STEP-KTODER框架,将代码步骤定义为多函数程序的模块级函数,结合函数级过程监督与程序结果级反馈,在多个代码基准测试中优于KTO、DPO等方法,且验证了执行标签的重要性。

Comments 20 pages, 8 figures, 14 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23311 2026-08-26 cs.CL 版本更新 57%

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

超越稳定性-探索困境:面向大语言模型策略优化的环境正则化方法

Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu

机构 * Alibaba Group(阿里巴巴集团) Beijing Normal University(北京师范大学) National University of Singapore(新加坡国立大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 该研究针对大语言模型策略优化的稳定性-探索困境,提出ERPO方法,通过查询KL散度正则化输入侧分布,接入现有策略优化流水线,在数学推理基准上实现了更强准确率与更稳定行为。

Comments Accepted to EMNLP 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14673 2026-08-26 cs.AI cs.GL quant-ph 版本更新 57%

Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument

AI生成数学证明的审计:量子并行重复中一个贪心条件引理的修正

Mikołaj Sienicki, Krzysztof Sienicki

机构 * Polish–Japanese Academy of Information Technology(波兰-日本信息技术学院) Chair of Theoretical Physics of Naturally Intelligent Systems (NIS)(自然智能系统理论物理研究所(NIS))

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本文审计OpenAI相关工作中AI生成的数学证明,发现其量子并行重复中贪心条件引理的证明存在极性错误,给出反例并提供局部修正证明,揭示AI论证可能隐藏互补事件的关键反转。

Comments 8 pages, 5 references. V2: 11 pages, 8 references. Substantially revised. The polarity error reported in v1 is withdrawn: it resulted from loss of an overbar during PDF text extraction, not from an error in the OpenAI manuscript. The revised paper retains and extends the independent human assessment of Chapter 6. No confirmed mathematical error was found in the examined argument

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22806 2026-08-25 cs.CL 新提交 57%

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

DIAG:面向数据高效数学偏好蒸馏的诊断式迭代对齐与生成

Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu

机构 * Tsinghua University(清华大学) Kyoto University(京都大学) Nanjing University(南京大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 DIAG是一种自适应调整训练分布的框架,通过诊断偏好对产出与生成针对性训练数据,提升数学推理任务中偏好监督的信息价值,在同等训练预算下增强模型性能。

Comments Accepted by EMNLP 2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22646 2026-08-25 cs.AI 新提交 57%

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

CAI-DLLM:面向扩散语言模型的感知收敛推理

Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 CAI-DLLM 是一种无需训练的扩散语言模型推理方法,通过利用第一步置信度调整解码策略,在多类任务上实现显著推理加速,同时保持较高准确率并降低能耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21496 2026-08-25 cs.LG math.ST stat.ML stat.TH 新提交 57%

The geometry of AI validation: Exact certification limits for iid best-of-N search

AI验证的几何:独立同分布N选最优搜索的精确认证极限

Ricardo Fitas

机构 * Technical University of Darmstadt(达姆施塔特工业大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 该研究针对独立同分布N选最优搜索推导了精确认证极限,提出双门审计规则,其在数学推理和代码选择的回顾性分析中可降低保留误差。

Comments 32 pages, 4 figures; includes Supporting Information. Code and processed data are available at https://github.com/rfitas-lab/geometry-of-ai-validation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14685 2026-08-25 cs.LG stat.ML 版本更新 57%

Rethinking Reverse KL as Adaptive Entropy Distillation

将反向KL重新思考为自适应熵蒸馏

Shizhen Li, Zhiyu Shen, Yuyin Lu, Yunhe Pang, Jielin Song, Yanghui Rao, Fu Lee Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 该研究针对知识蒸馏难以平衡忠实模仿与鲁棒生成的问题,提出自适应熵蒸馏(AED)方法,通过分解反向KL目标函数并利用教师熵动态校准模仿强度,在指令遵循与数学推理基准上表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00869 2026-08-25 cs.LG 版本更新 57%

Enhancing LLM Metacognition via Cognitive Pairwise Training

通过认知成对训练增强LLM元认知

Weitao Li, Hao Zhou, Xuanyu Lei, Fandong Meng, Yuanhang Liu, Jingyi Ren, Ante Wang, Xiaolong Wang, Yuanchi Zhang, Fuwen Luo, Guangwen Yang, Lin Gan, Weizhi Ma, Yang Liu

机构 * National Engineering Laboratory for Intelligent Information Processing, Academy of Mathematics and Physics, Chinese Academy of Sciences(智能信息处理国家工程实验室,中国科学院数学物理研究所) University of Science and Technology of China(中国科学技术大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 提出认知成对训练(CPT),通过成对比较推理轨迹来学习区分可靠与不可靠推理,从而提升LLM的推理与元认知权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13349 2026-08-25 cs.LG 版本更新 57%

When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration

少而精的潜在信息带来更好的中继:面向潜在多智能体大语言模型协作的信息保持压缩

Yiping Li, Zhiyu An, Wan Du

机构 * Department of Computer Science and Engineering(计算机科学与工程系) University of California, Merced(加州大学默塞德分校)

专题命中 数学推理 :reasoning(abstract);分类 cs.LG

AI总结 本文提出一种信息保持压缩方法,通过引入正交回填机制,在减少通信开销的同时保持信息完整性,提升了多智能体大语言模型协作性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20519 2026-08-24 physics.med-ph cs.AI 新提交 57%

An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy

用于磁共振引导放疗的扩散加权成像处理与解释一体化平台

Yunxiang Li, Yan Dai, Yen-Peng Liao, Jie Deng, Jill B De Vis, You Zhang

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 该研究开发了一体化网络平台,整合MR-Linac的DWI后处理与可追溯的RAG临床解释,经专家评分验证其在胶质母细胞瘤病例中具有高临床实用性。

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20314 2026-08-21 cs.AI 新提交 57%

MidTool: Mid-training Data Synthesis for Agentic Tool Use

MidTool:面向智能体工具使用的训练中期数据合成

Fengqing Jiang, Yite Wang, Boyi Liu, Zhaoyang Wang, Canwen Xu, Zhewei Yao, Radha Poovendran, Yuxiong He

机构 * University of Washington(华盛顿大学) Snowflake(雪花公司) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出MidTool流水线,构建面向智能体工具使用的训练中期数据,在MidTool-Mix上微调Qwen3模型,可显著提升工具使用相关下游任务性能,证明通用工具使用需专门训练中期调整。

Comments Data & Model: https://hf.co/collections/MidTool/midtool-release

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06127 2026-08-21 cs.AR cs.AI 版本更新 57%

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization

PrefixAgent:一种基于大语言模型的高效前缀加法器优化设计框架

Dongsheng Zuo, Jiadong Zhu, Yang Luo, Yuzhe Ma

机构 * HKUST(GZ)(香港科技大学(广州))

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究针对前缀加法器设计空间随位宽指数增长的优化难题,提出基于LLM的PrefixAgent框架,通过将任务拆分为两个子任务并利用等价图构建训练数据,在几乎所有配置下实现了更小的加法器面积,且在位宽更大时优势更显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19047 2026-08-20 cs.AI math.NT 新提交 57%

Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery

Eureka:面向科学发现的任务条件元智能体编排框架

Alizer Wong, Heng Cui, Yi Tan, Xiongchao Zhan, Liang Lin, Yuxiang Guo, Zhaorong Dai, Zixin Zeng, Wenyuan Li

机构 * ManXis(万熙科技(ManXis)) Guangdong University of Technology(广东工业大学) South China Normal University(华南师范大学) Shanghai Jiao Tong University(上海交通大学) Duke University(杜克大学) Hokkaido University(北海道大学)

专题命中 数学推理 :planning(abstract);分类 cs.AI

AI总结 该研究提出Eureka任务条件元智能体架构,通过动态编排完成科学发现任务,在递归任务、增量处理等实验中表现优异,还实现了数学领域的理论进展,证明科学智能体能力依赖匹配任务认知结构的架构。

Comments 62 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16977 2026-08-19 cs.AI math.CO 新提交 57%

The Problem Is the Problem: Towards Scalable Mathematical Discovery

问题本身才是问题:迈向可扩展的数学发现

Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck

机构 * Carnegie Mellon University(卡内基梅隆大学) Anysphere Co.(Anysphere公司)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 该研究提出人机协作的FAR流程,从文献中筛选数学问题并聚焦成果,在组合数学试点中筛选出77个待评审成果,验证了其数学发现的有效性。

Comments Code available at https://github.com/zeyu-zheng/FAR

详情

展开后加载摘要…

URL PDF HTML 收藏