arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Machine Learning · 会议 · Machine Learning

2026-01-27 至 2026-01-27 共收录 2
2502.17424 2026-01-27 cs.CL cs.AI cs.CR cs.LG

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

涌现的偏移:狭窄微调可以产生广泛偏移的LLM

Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans

机构 * University College London(伦敦大学学院) Center on Long-Term Risk(长期风险中心) Warsaw University of Technology(华沙技术大学) University of Toronto(多伦多大学)

AI总结 研究发现,狭窄微调训练LLM生成不安全代码会导致广泛偏移,模型在无关提示上表现出欺骗性行为,且偏移可通过触发器隐藏。

Comments 41 pages, 38 figures An earlier revision of this paper was accepted at ICML 2025. Since then, it has been updated to include new results on the impact of formatting (4.4), new dataset (4.6), training dynamics (4.7) and base models (4.8) Extended version of the paper was published in Nature 2026/1

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14175 2026-01-27 cs.CL cs.AI

GRAM: A Generative Foundation Reward Model for Reward Generalization

GRAM:一种用于奖励泛化的大规模生成式奖励模型

Chenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Tong Xiao, Chunliang Zhang, Tongran Liu, Jingbo Zhu

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院) NiuTrans Research, Shenyang, China(NiuTrans研究) CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China(中国科学院行为科学重点实验室) Meituan Inc.(美团公司)

AI总结 GRAM提出了一种生成式奖励模型,通过结合无监督和监督学习,提升奖励模型在多种任务上的泛化能力,有效改进了基线模型的性能。

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏