分叉提示的花园:用户如何在故事生成中探索叙事空间
The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation
- University of Colorado Boulder(科罗拉多大学博尔德分校)
- University of Georgia(佐治亚大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究利用真实用户与聊天机器人的对话日志,构建了WildStories和WildEdits数据集,提出编辑类型框架,分析用户通过迭代编辑提示词探索叙事空间的行为,并展示了基于该框架的自动排列用于故事生成基准测试。
AI中文摘要:
大型语言模型(LLMs)改变了人们与故事互动的方式。通过公开的聊天机器人日志,我们可以看到,当用户生成故事时,他们会迭代地编辑提示词以探索叙事可能性,调整角色、改变情节、切换虚构世界。作为聚合数据,这些提示词代表了大规模创意偏好的丰富痕迹。然而,故事生成评估基准依赖于静态的一次性提示词,无法捕捉这种探索行为。在这项工作中,我们研究了用户如何在真实环境中修改连续的故事提示词。利用自然发生的用户-聊天机器人对话数据集,我们构建了WildStories,一个包含275,635个故事生成提示词的样本(标注了故事格式、提示词组成部分和明确性),以及WildEdits,一个包含24,291个编辑树的集合,这些编辑树建模了用户如何迭代编辑基础故事提示词并探索分支故事可能性。从这些树中,我们开发了一个编辑类型框架,跨越四个方向(添加、删除、更改和扩展)与十四个目标(例如,情节、角色、类型)。然后,我们使用我们的数据集和这个框架来分析用户通过LLMs导航叙事空间的行为。最后,我们展示了基于该框架的自动排列如何用于故事生成基准测试。内容警告:本文涉及“野生”聊天机器人日志,这些日志通常包含有毒和性露骨主题。
英文摘要:
Large language models (LLMs) have changed the way people engage with stories. Drawing on public chatbot logs, we can see that when users generate stories, they iteratively edit their prompts to explore narrative possibilities, adjusting characters, redirecting plots, and swapping fictional universes. As aggregated data, these prompts represent rich traces of creative preference at scale. Yet story generation evaluation benchmarks rely on static, one-shot prompts that cannot capture this exploratory behavior. In this work, we study how users revise consecutive story prompts in the wild. Using a dataset of naturally occurring user-chatbot conversations, we construct WildStories, a sample of 275,635 story generation prompts (labeled with story format, prompt components, and explicitness), and WildEdits, a collection of 24,291 edit trees that model how users iteratively edit base story prompts and explore branching story possibilities. From these trees we develop a framework of edit types crossing four directions (adding, removing, changing, and extending) with fourteen targets (e.g., plot, character, genre). We then use our datasets and this framework to analyze user behavior in navigating narrative space via LLMs. Finally, we show how automated permutations based on the framework can be used for story generation benchmarking. Content Warning: This paper works with "wild" chatbot logs, which often include toxic and sexually explicit themes.