Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
Andrii Balashov
专题命中
后训练与偏好优化
:large language model(title,abstract);language model(title,abstract);pretraining(abstract);分类 cs.AI、cs.LG
CommentsThe manuscript has been withdrawn due to significant overlap in methodology and results with a prior work (arXiv:2505.11711) that we were not aware of at the time of submission. To maintain academic integrity and avoid redundancy in the literature, we have chosen to withdraw this version
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
Anas Mohamed, Azal Ahmad Khan, Xinran Wang, Ahmad Faraz Khan, Shuwen Ge, Saman Bahzad Khan, Ayaan Ahmad, Ali Anwar
机构
*
University of Minnesota(明尼苏达大学)
;
Virginia Tech(弗吉尼亚理工大学)
;
Xi’an University of Technology(西安理工大学)
;
Lahore University of Management Sciences(拉合尔管理科学大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)