MUSE:一个理论驾驭的故事引擎,用于氛围叙事化
MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing
浏览论文内容
中文总结 AI 辅助
MUSE是一个基于麦基故事理论、通过智能体框架和上下文工程将自然语言要求转化为高质量故事的故事引擎,在多个基准上显著提升生成质量。
中文摘要 AI 辅助
大语言模型(LLM)能够生成流畅的散文。故事质量取决于情节、人物和语言方面的决策如何在规划、起草和修改过程中协同作用。引导这些决策面临两个瓶颈:故事指导的质量及其持续使用。我们将氛围叙事化(Vibe Narrativizing)定义为将自然语言写作要求转化为完整故事的任务,并提出了MUSE——一个理论驾驭的故事引擎。MUSE将故事知识组织为针对特定决策的指导,并将这些决策延续到后续的创作工作中。知识工程通过规则原子化、语义整合和机制抽象来发展罗伯特·麦基(Robert McKee)的故事理论;单一事实来源和分层披露组织所产生的指导。典型示例补充了依赖上下文和审美判断的原则。一个智能体框架通过保留故事决策的中间交付物来组织设计、角色表演、场景构图和修改。上下文工程为每个角色提供相关的指导和决策,而杰作语料库提供灵感和散文参考。一个工作示例跟随一个请求的对象从其主题角色到角色的高潮行动。在四个基础模型上,MUSE在WritingBench上比零样本生成提高了1.6-4.8分,并在三个模型上将LongStoryEval提高了十分以上。ConStory-Bench的一致性错误密度在所有四个模型中保持在低个位数,低于三个复现的故事系统基线。组件消融实验发现最大的质量贡献在于结构设计,角色路径中的特定声音效果,以及修改中的进一步收益。代码可在以下网址获取:此https URL。
英文摘要
LLMs have been able to generate fluent prose, but high-quality stories also require coordinated decisions about plot, character, and language across planning, drafting, and revision. We formulate Vibe Narrativizing as turning natural-language writing requirements into a finished story. MUSE, a Theory-Harnessed Story Engine, addresses two bottlenecks: rule quality and sustained rule realization. Story theory supplies the rules, and a practical agent harness puts them to work. Knowledge engineering organizes Robert McKee's theory through rule atomization, semantic consolidation, mechanism abstraction, a single source of truth, and layered disclosure; typical examples clarify judgments that depend on context and aesthetic purpose. The harness preserves story decisions in intermediate deliverables across design, character performance, scene composition, and revision. Context engineering supplies each role with relevant guidance and decisions; a masterwork corpus provides inspiration and prose references. A worked example follows a requested object from its thematic role to climactic actions. Across four base models, MUSE improves WritingBench by 1.1 to 6.2 points over zero-shot generation; it is the only multi-stage system in our comparison to do so. It also raises LongStoryEval by more than ten points on three of the four models. ConStory-Bench consistency error density remains in the low single digits for all four models, below every reproduced story-system baseline on three of the four models. Ablations locate the largest quality contribution in structural design, voice-specific effects in the character path, and further gains in revision.
发表机构
- Northeastern University(东北大学)
- OranAI
- OranAI Ltd.(OranAI有限公司)
机构由 AI 辅助整理,请以论文原文为准。