CommentsA technical report from MiniMax. The authors are listed in alphabetical order. We open-source our MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1
CommentsAccepted at the Conference on Language Modeling (COLM) 2026. 45 pages, including appendices; 24 figures and 12 tables. Code: https://github.com/withmartian/mi-cot
The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
Raul Cavalcante Dinardi, Bruno Yamamoto, Anna Helena Reali Costa, Artur Jordao
机构
*
Instituto de Matemática, Estatística e Ciência da Computação, Universidade de São Paulo(圣保罗大学数学、统计与计算机科学研究所)
;
Escola Politécnica, Universidade de São Paulo(圣保罗大学理工学院)
Efficient Process Reward Modeling via Contrastive Mutual Information
通过对比互信息实现高效的进程奖励建模
Nakyung Lee, Sangwoo Hong, Jungwoo Lee
机构
*
Department of Electrical and Computer Engineering, Seoul National University(首尔大学电气与计算机工程系)
;
Department of Computer Science and Engineering, Konkuk University(建国大学计算机科学与工程系)
MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action
MPCoT: 奖励引导的多路径潜在推理用于测试时可扩展的视觉-语言-动作
Boyang Zhang, Lianlei Shan
机构
*
Department of Electrical and Computer Engineering, Boston University(波士顿大学电气与计算机工程系)
;
Department of Computer Science, Tsinghua University(清华大学计算机系)