发表机构
Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究引入Rushes数据集,通过游戏界面收集用户与AI叙事交互数据。核心方法是分析用户选择模式,以量化非随机偏好。主要贡献是定位其为多元对齐诊断基准,揭示现有模型在捕捉个性化参与信号上的不足,助力相关研究。
AI 中文摘要
我们引入了Rushes,这是一个用于研究交互式叙事环境中人类显性参与偏好的数据集和基准。Rushes通过游戏界面收集,用户与人工智能生成的分支叙事进行交互,并在每个决策点从一个小的、明确的候选集中选择一个选项。每次交互都会记录完整的候选集、用户的选择以及不断演变的叙事背景,生成带有持久用户级标识符的时间顺序轨迹。Rushes包含来自六款游戏中8167个独特用户的44226个决策事件,捕捉连续的、个性化的参与行为而非静态判断。我们表明用户选择呈现出结构化、非随机模式,通过相对于均匀基线的低选择熵来量化。我们将Rushes定位为多元对齐的诊断基准,并展示了一个强大的参与差距:包括GPT-5在内的最先进的语言模型未能超越简单基线。经典矩阵分解(SVD)能捕捉到可测量的个性化信号(37.7%),前沿语言模型(34.23%)在事件级选择预测上甚至难以与流行度基线(36.4%)相匹配。这一差距表明,现代基于人类反馈的强化学习中使用的单一总体目标似乎不足以捕捉异质的、依赖上下文的参与信号。因此,即使是能力很强的模型也默认遵循多数偏好,而不是适应个体轨迹。我们发布Rushes以支持对生成系统中的多元对齐和顺序决策的研究。平台和数据集的完整代码将在此处提供:此https URL
英文摘要
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching narratives and select one choice from a small, explicit candidate set at each decision point. Each interaction logs the full candidate set, the user's choice, and the evolving narrative context, yielding time-ordered trajectories with persistent user-level identifiers. Rushes contains 44,226 decision events from 8,167 unique users across six games, capturing sequential, personalized engagement behavior rather than static judgments. We show that user choices exhibit structured, non-random patterns, quantified by a low choice entropy relative to a uniform baseline. We position Rushes as a diagnostic benchmark for pluralistic alignment and demonstrate a robust Engagement Gap: state-of-the-art LLMs, including GPT-5, fail to outperform simple baselines. While classical Matrix Factorization (SVD) captures measurable personalized signal (37.7%), frontier LLMs (34.23%) struggle to even match the Popularity Baseline (36.4%) on event-level choice prediction. This gap suggests that single, population-level objectives, like those used in modern RLHF, appear insufficient to capture heterogeneous, context-dependent engagement signals. As a result, even highly capable models default to majority preferences rather than adapting to individual trajectories. We release Rushes to support research into pluralistic alignment and sequential decision-making in generative systems. The full code for the platform and dataset will be available here: https://github.com/microsoft/rushes