MathScale: Scaling Instruction Tuning for Mathematical Reasoning
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Work in progress
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Work in progress
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 12 pages, 5 figures
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 18 pages, 9 figures, 12 tables
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 116 pages, 120 figures. Accepted to ICLR 2024
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments AAAI 2024
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at the Findings of ACL 2023
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2023. The repository is available at https://github.com/lupantech/dl4math
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Project Webpage and Code: https://composable-models.github.io/llm_debate/
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL 2022
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments The paper has been accepted to the ACL-IJCNLP 2021 conference
ProgramTab:通过编程范式提升大语言模型的表格推理能力
机构 * Big Data and AI Platform Department, Tencent(腾讯大数据与人工智能平台部) ; Institute of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 研究基于大语言模型的表格推理问题,提出ProgramTab框架,指导LLMs用Python代码预处理表格数据并进行关键内容提取,实验证明该框架能有效处理表格推理任务,性能优于基于LLM的基线。
Comments Large Language Models, Table Reasoning, In-context Learning
学习如何使用工具,而非仅仅何时:基于模式的工具集成推理
机构 * University of Georgia(佐治亚大学) ; University of Maryland, Baltimore County(马里兰大学巴尔的摩分校) ; The University of Hong Kong(香港大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出一种两阶段框架,通过构建代码能力并对齐模式选择与教师偏好,提升工具集成推理的代码使用和准确性,实验显示在数学数据集上显著提升。
Journal ref The 5th Workshop on Mathematical Reasoning and AI at NeurIPS 2025
通过混合训练缓解数学推理微调中的灾难性遗忘
机构 * Department of Computer Science(计算机科学系) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI总结 本文提出混合训练策略,通过交错数学和NLI任务来缓解微调中的灾难性遗忘,同时保持数学性能和NLI准确性。
Comments 11 pages, 2 figures. Code available at https://github.com/johngrahamreynolds/mathematical_catastrophe_mitigation. Models available at https://huggingface.co/collections/MarioBarbeque/catastrophic-forgetting-in-mathematical-reasoning
机构 * MIT(麻省理工学院) ; Comenius University in Bratislava(布拉迪斯拉发孔院大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments Accepted to The 5th Workshop on Mathematical Reasoning and AI at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025); 25 pages, 14 figures, 8 tables; Code available at https://github.com/NaiveNeuron/FractalBench
机构 * UC San Diego(加州大学圣地亚哥分校) ; Capital One(Capital One公司)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments NeurIPS 2025 Workshop on Efficient Reasoning
机构 * IBM Research – Zurich(IBM瑞士研究中心) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments Accepted at the 5th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2025
机构 * KAIST AI(韩国科学技术院人工智能研究所) ; KRAFTON(KRAFTON公司) ; UC Berkeley(伯克利大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments Accepted into Test-time Scaling and Reasoning Models (SCALR) workshop at COLM 2025. 28 pages
机构 * University of California Los Angeles(加州大学洛杉矶分校) ; School of Artificial Intelligence Chinese Academy of Sciences(中国科学院人工智能学院) ; Microsoft(微软) ; Tsinghua University(清华大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG
Comments Reinforcement Learning; Large Language Models; LLM Reasoning
是解码格式,而非扰动:审计视觉语言模型测试时扩展的基于一致性的选择
专题命中 数学推理 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.AI
AI总结 该研究发现视觉语言模型测试时的选择效果由解码格式而非扰动决定,提出的扰动基选择(Pgs)在控制格式后无显著收益,说明其无法作为有效选择信号。
HoT: 用于从输入中引用支持事实的高亮推理链
机构 * Auburn University(亚伯拉罕大学) ; University of Alberta(阿尔伯塔大学)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);logical reasoning(abstract)
AI总结 本文提出HoT技术,通过XML标签高亮输入中的关键事实,减少LLM的幻觉问题,提升22项任务的准确性,并帮助人类更高效地验证结果。
Formula-One Prompting:一种可组合的方程优先前缀用于应用数学
机构 * SCB DataX, SCBX Group(SCB数据X,SCBX集团)
专题命中 数学推理 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.CL
AI总结 提出公式提示(FP)和Formula-One提示(F-1),通过先形式化问题中的控制方程再求解,在多个应用数学基准上优于思维链和程序思维提示,平均提升5.76和8.42个百分点。
大型语言模型在语义保持变异下的理解鲁棒性如何?
机构 * Department of Computer Science University of Oxford(计算机科学系牛津大学)
专题命中 数学推理 :reasoning(summary_cn,abstract);分类 cs.AI
AI总结 本文评估了大型语言模型在Python程序上的推理能力,通过五种语义保持的代码变异测试,发现模型在语义变换下表现出显著的脆弱性,正确预测基于 flawed reasoning 的比例在10%-50%之间,且预测稳定性差。
Comments 17 pages, 5 tables, 1 figure
令牌级策略优化:通过序列级似然将群体级奖励与令牌级聚合联系起来
机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) ; Key Laboratory of Symbolic Computation and Knowledge Engineering of MOE, Jilin University(教育部符号计算与知识工程重点实验室,吉林大学) ; Baidu Inc.(百度公司) ; Tsinghua University(清华大学) ; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室,计算技术研究所)
专题命中 数学推理 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.CL
AI总结 本文提出TEPO,通过序列级似然将群体奖励与单个令牌联系,并引入KL散度掩码约束以提升训练稳定性,实验显示其在数学推理任务中表现优异且收敛时间减少50%。
机构 * PhysicsWallah(物理墙)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);math reasoning(abstract)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);logical reasoning(abstract)
Comments Published at ICLR 2024
用于物理推理的解耦物理建模与执行
机构 * University of Pennsylvania(宾夕法尼亚大学) ; William & Mary(威廉玛丽学院) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Amazon(亚马逊公司) ; Stevens Institute of Technology(史蒂文斯理工学院) ; Northwestern University(西北大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI总结 该研究提出解耦物理建模与执行的统一框架,通过两阶段后训练策略提升小型大语言模型物理推理性能,在多个基准上实现约3%的平均性能提升。
Olapa-MCoT:增强大语言模型的中文数学推理能力
机构 * Tencent(腾讯) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本研究针对Llama-2-13B中文数学推理能力弱的问题,提出Olapa-MCoT方法,通过SimRRHF和IDRL提升模型性能,使中文推理准确率升36%,且可适配任意主流LLMs。
Comments 10 pages, 1 figures
为什么自蒸馏(有时)会降低大语言模型的推理能力?
机构 * Microsoft Research(微软研究院) ; KAIST(韩国成均馆大学) ; Seoul National University(首尔国立大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI总结 本文研究了自蒸馏在数学推理中降低大语言模型推理能力的原因,发现其通过抑制模型在推理过程中的不确定性表达,导致在未见过的问题上表现下降,强调了适当表达不确定性对鲁棒推理的重要性。
Comments Accepted to COLM 2026. Code is available at https://github.com/beanie00/self-distillation-analysis