Entropy-Regularized Process Reward Model
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Toronto(多伦多大学) ; NVIDIA(英伟达) ; Princeton University(普林斯顿大学) ; Salesforce Research(Salesforce研究)
Comments Upate TMLR version
期刊&会议
Transactions on Machine Learning Research · 期刊 · Machine Learning
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Toronto(多伦多大学) ; NVIDIA(英伟达) ; Princeton University(普林斯顿大学) ; Salesforce Research(Salesforce研究)
Comments Upate TMLR version
机构 * University of Helsinki(赫尔辛基大学)
Comments Final version published at TMLR
机构 * Department of Computer Science ETH Zurich(计算机科学系,苏黎世联邦理工学院) ; Research Center Trustworthy Data Science and Security of the University Alliance Ruhr(可信数据科学与安全的鲁尔大学联盟研究中心) ; Department of Statistics TU Dortmund University(统计学系,多特蒙德技术大学)
Comments Published in Transactions on Machine Learning Research (TMLR) https://openreview.net/forum?id=Bt0zdsnWYc