LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
揭开真相的面纱:秩减少后主权重的浮现使以推理为导向的监督微调得以提升
Zihang Liu, Tianyu Pang, Oleg Balabanov, Chaoqun Yang, Tianjin Huang, Lu Yin, Yaoqing Yang, Shiwei Liu
机构
*
University of California, Berkeley, CA, USA(加州大学伯克利分校)
;
International Computer Science Institute, CA, USA(国际计算机科学研究所)
;
Lawrence Berkeley National Laboratory, CA, USA(伯克利国家实验室)
;
Dartmouth College, NH, USA(达特茅斯学院)
;
University of Exeter, Exeter, UK(埃克塞特大学)
;
University of Oxford, Oxford, UK(牛津大学)
;
University of Surrey, Guildford, UK(萨里大学)
;
Tsinghua University, China(清华大学)
;
Eindhoven University of Technology, the Netherlands(埃因霍温理工大学)
机构
*
Tencent AI Seattle Lab(腾讯AI西雅图实验室)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of Maryland, College Park(马里兰大学科廷分校)
;
The University of Texas at Dallas(德克萨斯大学达拉斯分校)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
是什么使低比特量化感知训练在推理大语言模型中有效?一项系统研究
Keyu Lv, Manyi Zhang, Xiaobo Xia, Jingchen Ni, Shannan Yan, Xianzhi Yu, Lu Hou, Chun Yuan, Haoli Bai
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Huawei Technologies(华为技术有限公司)
;
National University of Singapore(新加坡国立大学)
C$^2$GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
C$^2$GSPG:基于置信度校准的群体序列策略梯度方法,用于自感知推理
Haotian Liu, Shuo Wang, Hongteng Xu
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高陵人工智能学院)
;
Tsinghua University(清华大学)
;
Beijing Key Laboratory of Research on Large Models(北京大模型研究关键实验室)
;
Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
套利:通过优势感知投机实现高效推理
Monishwaran Maheswaran, Rishabh Tiwari, Yuezhou Hu, Kerem Dilmen, Coleman Hooper, Haocheng Xi, Nicholas Lee, Mehrdad Farajtabar, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
对LLMs在数学推理中鲁棒性的调查:通过高级数学问题的数学等价转换进行基准测试
Yuren Hao, Xiang Wan, ChengXiang Zhai
机构
*
Department of Computer Science University of Illinois Urbana–Champaign(计算机科学系伊利诺伊大学厄巴纳-香槟分校)
;
Department of Computer Science Stanford University(计算机科学系斯坦福大学)
机构
*
Department of Electrical & Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)
;
AI Lab, Princeton University(普林斯顿大学人工智能实验室)
;
Department of Computer Science & Engineering, University of Michigan(密歇根大学计算机科学与工程系)
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
Shanghai Qi Zhi Institute(上海启智研究所)
;
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)