GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
GeoMathCode: 理解几何问题求解中交织的数学-代码推理
Yingji Zhang, Yong Dai, André Freitas
机构
*
Idiap Research Institute(Idiap研究 institute)
;
X-Humanoid
;
Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系)
;
Cancer Biomarker Centre, CRUK Manchester Institute(癌症生物标志物中心,CRUK曼彻斯特研究所)
MUR: Momentum Uncertainty guided Reasoning for Large Language Models
MUR: 面向大型语言模型的动量不确定性引导推理
Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Luu Anh Tuan, Haiteng Zhao, Qika Lin, Jun Liu
机构
*
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
Ministry of Education Key Laboratory of Intelligent Networks and Network Security(教育部智能网络与网络安全重点实验室)
;
Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室)
;
Nanyang Technological University(南洋理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
VinUniversity(Vin大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Northwestern University(西北大学)
;
University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Pittsburgh(匹兹堡大学)
Unveiling the Reasoning Process of Large Language Models
揭示大型语言模型的推理过程
Junjie Zhang, Zhen Shen, Xisong Dong, Gang Xiong
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training
隐式压缩正则化:通过内部更短分布实现简洁推理(在强化学习后训练)
Chen Wang, Hexuan Deng, Yining Zhang, Yuchen Zhang, Jionghao Bai, Zhaochun Li, Ge Lan, Yue Wang
机构
*
College of Software, Nankai University(南开大学软件学院)
;
Zhongguancun Academy(中关村学院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
East China Normal University(华东师范大学)
;
Zhejiang University(浙江大学)
;
Beijing Institute of Technology(北京理工大学)
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
用推理树调度你的LLM强化学习
Hong Wang, Zhezheng Hao, Jian Luo, Chenxing Wei, Yao Shu, Lei Liu, Qiang Lin, Hande Dong, Jiawei Chen
机构
*
Cranberry-Lemon University(Cranberry-Lemon 大学)
;
University of the Witwatersrand(独立研究者)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
机构
*
Shanghai University of Finance and Economics(上海金融学院)
;
Ant Group(蚂蚁集团)
;
Southern University of Science and Technology(南方科技大学)
;
MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(计算与经济学交叉学科研究教育部重点实验室)