发表机构
University of Macau; Westlake University; Harbin Institute of Technology; University of Cambridge; University of Aberdeen(澳门大学; 西湖大学; 哈尔滨工业大学; 剑桥大学; 阿伯丁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体在交互式叙事中难以维持长期一致性的问题,构建了NCP-Bench基准测试,发现现有顶尖LLM的承诺保持率较低,存在显著的逻辑冲突问题。
AI 中文摘要
大语言模型(LLM)的快速发展正通过支持开放式、流畅的交互式叙事,革新游戏领域的人工智能。然而,现有研究大多忽视了一个关键挑战:在不受约束的用户干预下维持长期的逻辑一致性和叙事完整性。为解决这一问题,我们将该挑战定义为叙事承诺保持(Narrative Commitment Preservation, NCP),并以交互式叙事作为测试平台。我们推出NCP-Bench,这是一个由100个源自电影梗概的叙事环境组成的基准测试。每个环境包含结构化的叙事规范(轨迹、承诺和初始事实),我们可在玩家智能体与旁白智能体的交互过程中自动检查这些规范。对现有最先进LLM的实验显示,存在显著的长期一致性差距:高语言质量并不保证承诺保持;即便是强大的模型在对抗性干预下也频繁生成逻辑冲突的内容,表现最佳的模型(GPT-5.2)在20轮交互后的承诺保持率仅为42%,各模型的事实冲突率在40%至68%之间,且仅有少数运行能在100轮的限制内满足所有成就承诺。
英文摘要
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses. Each environment includes a structured narrative specification (trajectory, commitments, and initial facts) that we can automatically check throughout the interaction between the player agent and the narrator agent. Experiments across state-of-the-art LLMs reveal a substantial long-horizon consistency gap: high linguistic quality does not guarantee commitment preservation; even strong models frequently generate logically conflicting content under adversarial interventions, with the best-performing model (GPT-5.2) achieving only 42% survival rate after 20 turns and fact conflict rates ranging from 40% to 68% across models, and only isolated runs satisfying all achievement commitments within the 100-turn limit.
CommentsAccepted by ICML 2026