arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-17 至 2025-12-17 共收录 5 信号源:cs.CL, cs.AI, cs.LG

1. 数学推理 5 篇

2511.01170 2025-12-17 cs.AI 83%

DART: Difficulty-Adaptive Reasoning Truncation for Efficient Large Language Models

DART: 为高效大语言模型的难度自适应推理截断

Ruofan Zhang, Bin Xia, Zhen Cheng, Cairen Jian, Minglun Yang, Ngai Wong, Yuan Cheng

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) CSE Department, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Research and Development Department, SIMMIR Tech(Simmir科技研发部) School of Information Science and Technology, Xiamen University Tan Kah Kee College(厦门大学信息科学与技术学院) The University of Hong Kong(香港大学)

专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 DART通过自适应调整推理长度提升大语言模型效率,实现81.2%的推理截断和5.33倍的计算加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13706 2025-12-17 cs.LG cs.CL 82%

Mitigating Catastrophic Forgetting in Mathematical Reasoning Finetuning through Mixed Training

通过混合训练缓解数学推理微调中的灾难性遗忘

John Graham Reynolds

机构 * Department of Computer Science(计算机科学系) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出混合训练策略,通过交错数学和NLI任务来缓解微调中的灾难性遗忘,同时保持数学性能和NLI准确性。

Comments 11 pages, 2 figures. Code available at https://github.com/johngrahamreynolds/mathematical_catastrophe_mitigation. Models available at https://huggingface.co/collections/MarioBarbeque/catastrophic-forgetting-in-mathematical-reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13978 2025-12-17 cs.AI 79%

Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms

评估前沿大语言模型在博士级数学推理中的表现:一个关于随机算法理论教材的基准测试

Yang Cao, Yubin Chen, Xuyang Guo, Zhao Song, Song Yue, Jiahao Zhang, Jiale Zhao

专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI

AI总结 本文评估了前沿大语言模型在博士级数学推理中的表现,通过随机算法教材的基准测试,发现顶级模型在准确性上表现优异,但可靠性存在显著差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14237 2025-12-17 cs.CL cs.LG 79%

Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets

向上爬,记忆下降:低成本微调与侧边网络

Estelle Zheng, Nathan Cerisara, Sébastien Warichet, Emmanuel Helbert, Christophe Cerisara

专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.LG

AI总结 LST通过轻量级侧边网络实现高效微调,比QLoRA更节省内存,同时在多个任务中保持竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10966 2025-12-17 physics.ed-ph 50%

Can Large Language Models Correctly Interpret Equations with Errors?

大型语言模型能否正确解释有误的方程?

Lachlan McGinness, Peter Baumgartner

专题命中 数学推理 :reasoning(abstract)

AI总结 本文研究了大型语言模型在解析学生输入中存在错误的方程时的准确性,发现开源模型无法达到预期效果,并提出未来改进方向。

Comments Published in Physics Review Physics Education Research here: https://journals.aps.org/prper/abstract/10.1103/v8f8-s11v

Journal ref Phys. Rev. Phys. Educ. Res. 21, 020155 Published 15 December, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏