GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) ; School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院(iOPEN),西北工业大学) ; HuaWei Technologies Co., Ltd.(华为技术有限公司) ; The University of Hong Kong(香港大学)