arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增量指令传递对语言模型创意写作的影响

The Effects of Incremental Instruction Delivery on Language-Model Creative Writing

Anshuman Singh, Abrar Eyasir, Haseeb Yaqoob, John Manavalan

arXiv 2609.33738首次发表:更新:

发表机构

SGT UNIVERSITY; University of Dhaka; NED University of Engineering and Technology; Metea Valley High School(SGT大学; 达卡大学; NED工程技术大学; 米蒂亚谷高中)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过160个创意写作任务和960个匹配对,发现增量指令交付降低约束遵循度并损害结构连贯性,提出创意完整性度量,表明交互式系统需评估要求的连贯整合。

AI 中文摘要

大型语言模型越来越多地被用作交互式写作工具,用户在多轮对话中逐步发展故事、修改想法并引入新要求,而不是事先提供完整的简要说明。然而,关于多轮指令退化的多数证据来自具有客观可验证结果的任务,这使得增量交互是否会以显式要求检查无法捕捉的方式损害创意产物仍不清楚。我们使用六个体裁的160个由人类撰写的创意写作任务来研究这一问题,将每个预期规范要么一次性提供,要么在5至9轮中逐步提供给六个不同的开放权重模型家族,产生960个匹配对。渐进式交付降低了显式约束遵循度,并在结构/连贯性方面产生了最大的写作质量退化。在具有相同观察遵循度的输出中,结构差距仍然存在,这表明仅凭测量的要求损失并不能解释观察到的结构差异。我们将创意完整性定义为联合遵循度和叙事结构的紧凑度量;在增量交付下,模型保留了完整创意完整性的71.2%(95%置信区间[68.2%, 74.3%])。一项由三名评分者进行的50个匹配对的人类研究独立恢复了完整交付在结构/连贯性、技巧和体裁有效性方面的优势,而自动化分数与汇总的人类评分保持正相关。这些发现表明,交互式创意写作系统不仅应评估要求是否在对话中存活,还应评估不断演变的要求是否被连贯地整合到最终产物中。我们的数据集、基准和源代码可在以下网址获取:此https URL

英文摘要

Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with objectively verifiable outcomes, leaving unclear whether incremental interaction harms creative artifacts in ways that explicit requirement checks cannot capture. We study this question using 160 human-authored creative-writing tasks across six genres, presenting each intended specification either upfront or progressively over 5-9 turns to six distinct open-weight model families, yielding 960 matched pairs. Progressive delivery reduces explicit constraint adherence and produces its largest writing-quality degradation in structure/coherence. The structural gap persists among outputs with equal observed adherence, suggesting that measured requirement loss alone does not explain the observed structural difference. We define Creative Integrity as a compact measure of joint adherence and narrative structure; under incremental delivery, models retain 71.2% of FULL Creative Integrity (95% CI [68.2%, 74.3%]). A three-rater human study over 50 matched pairs independently recovers FULL advantages in structure/coherence, craft, and genre effectiveness, while automated scores remain positively associated with aggregated human ratings. These findings show that interactive creative-writing systems should be evaluated not only on whether requirements survive conversation, but also on whether evolving requirements remain coherently integrated into the final artifact. Our dataset, benchmarks, and source code are available at: https://github.com/solusops/SISTER-2026-Team19

Comments18 pages, 4 figures, 13 tables. Code and data available at https://github.com/solusops/SISTER-2026-Team19

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑