Learning from Mistakes: Negative Reasoning Samples Enhance Out-of-Domain Generalization
从错误中学习:负推理样本增强领域外泛化
机构 * CAS Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(中国科学院人工智能安全重点实验室,计算技术研究所,中国科学院,北京,中国) ; Harbin Institute of Technology, Harbin, China(哈尔滨工业大学,哈尔滨,中国) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) ; Tsinghua University, Beijing, China(清华大学,北京,中国)
专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL
AI总结 通过引入负例推理样本,GLOW方法有效提升了大语言模型在领域外任务中的泛化能力,通过调节损失下降和提升策略熵来减少过拟合并促进探索。
Comments Code and data are available at https://github.com/Eureka-Maggie/GLOW