arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CreativeInstruct:规模化教导大型语言模型(LLM)平衡质量、创造性与多样性

CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity

Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin

arXiv 2608.07460首次发表:更新:

发表机构

Columbia; UNC Chapel Hill; University of Texas at Austin(哥伦比亚大学; 北卡罗来纳大学教堂山分校; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CreativeInstruct指令调优方法,通过注入[StartCreativity]跨度平衡LLM的质量、创造性与多样性,其在叙事生成中表现更优,还能提升强化学习任务性能。

AI 中文摘要

尽管后训练可提升大型语言模型(LLM)的能力,但通常会降低其输出的多样性和创造性,对明确需要创造性的任务(如故事生成)以及隐含需要创造性的任务(如强化学习(RL))产生负面影响。我们提出CreativeInstruct,这是一种可规模化的指令调优方法,通过学习注入特殊的[StartCreativity]跨度来引导生成偏向创造性,从而教导LLM平衡类似基础模型的创造性生成与后训练模型的质量。此外,我们引入了一种基于图编辑距离的结构多样性度量,该度量可捕捉纯词汇和语义度量遗漏的叙事层面变化。在叙事生成任务中,CreativeInstruct的多样性与多模型基线及其输出的蒸馏变体相当或更优,且未牺牲质量,推理时也无需多个模型。我们的人工评估结果与此一致,标注者认为CreativeInstruct的生成内容比后训练LLM的生成内容更具创造性的比例达70.3%。我们还表明,创造性模型可作为强化学习的基础:将GRPO应用于CreativeInstruct检查点,与应用于后训练检查点的相同训练相比,在AMC上提升约4%,在MATH上提升约5个百分点。

英文摘要

While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.

CommentsCode: https://github.com/ananya-sahu/CreativeInstruct

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑