Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) ; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL