LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
机构 * UNC Chapel Hill(杜克大学夏洛特分校) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 测试时计算 :reasoning(abstract);math reasoning(abstract);分类 cs.CL、cs.LG
Comments NeurIPS 2025 camera-ready. First two authors contributed equally. Code: https://github.com/duykhuongnguyen/LASeR-MAB