Provably Efficient Online RLHF with One-Pass Reward Modeling
机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术实验室,南京大学) ; School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)
Comments NeurIPS 2025; The first two authors contributed equally