arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

ByteDance(字节跳动)

2026-07-09 至 2026-07-09 共收录 2
2607.06987 2026-07-09 cs.LG 新提交

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

UP:用于打破探索-稳定性困境的无界正不对称优化

Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin

机构 * ByteDance Seed(字节跳动种子) Michigan State University(密歇根州立大学)

AI总结 研究针对强化学习探索-稳定性困境,提出无界正不对称优化(UP)方法,通过特殊设计重组优化过程,在多种算法、模型架构及训练模式下增强探索能力,实现卓越推理精度,是通用即插即用的强化学习训练增强方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05394 2026-07-09 cs.LG cs.AI cs.CL 新提交

Weak-to-Strong Generalization via Direct On-Policy Distillation

通过直接在线策略蒸馏实现弱到强泛化

Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu, Zhilong Zhang, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

机构 * SIA-Lab of Tsinghua AIR and ByteDance Seed(清华-字节跳动联合研究中心SIA实验室) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Peking University(北京大学)

AI总结 针对可验证奖励RL在大模型上成本高的问题,提出Direct-OPD方法迁移小模型RL的策略偏移,无需在大模型上跑RL即可实现性能提升。

Comments Project Page: https://bytedtsinghua-sia.github.io/Direct-OPD/

详情

展开后加载摘要…

URL PDF HTML 收藏