arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于强化学习的故事建设性反馈生成

Generating Constructive Feedback on Stories via Reinforcement Learning

Maja Stahl, Timon Ziegenbein, Henning Wachsmuth

arXiv 2609.04824首次发表:更新:

发表机构

Leibniz University Hannover; L3S Research Center(汉诺威莱布尼茨大学; L3S研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM生成反馈泛化、缺乏可操作性的问题,提出基于GRPO强化学习的方法,在三个故事语料库评估中优于Gemini等模型,证实可操作建议是反馈建设性的核心驱动。

AI 中文摘要

建设性反馈对创意作家提升叙事能力至关重要,由于获取人类专家反馈往往成本高且耗时,大语言模型(LLMs)作为自动写作助手提供了一种可扩展且高效的替代方案。尽管LLMs具有潜力,但研究表明其生成的反馈往往泛化性强、缺乏可操作性,且无法识别最关键的写作问题。为解决这些局限,本文提出一种强化学习方法,无需真实反馈即可引导LLMs生成建设性反馈。我们采用组相对策略优化(GRPO)训练模型,该模型带有新型多组件奖励函数,旨在实现建设性:优先选择针对故事量身定制、有助于提升故事质量且解决最关键写作问题的反馈。在三个故事语料库上的自动与人工评估中,我们的方法优于包括Gemini在内的最先进LLMs及竞争性基线,且发现提供可操作建议是反馈建设性的主要驱动因素。

英文摘要

Constructive feedback is crucial for creative writers to refine their storytelling abilities. Since receiving feedback from human experts is often costly and time-intensive, large language models (LLMs) offer a scalable and efficient alternative as automatic writing assistants. Despite their potential, research indicates that LLM-generated feedback is often generic, lacks actionability, and fails to identify which writing issue is most critical. To address these limitations, we present a reinforcement learning approach that steers LLMs to generate constructive feedback without the need for ground-truth feedback. We train our model using group relative policy optimization (GRPO) with a novel multi-component reward function aiming at constructiveness: it prioritizes feedback that is uniquely tailored to the story, helps to improve story quality, and addresses the most critical writing issue. In automatic and human evaluation across three story corpora, our approach outperforms state-of-the-art LLMs (including Gemini) and competitive baselines. We find that providing actionable suggestions is the main driver of feedback constructiveness.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑