AI 中文总结
该研究构建了带真实特质的合成玩家基准,提出机会感知决策时刻表示,对比了LLM等方法的推断效果,将推断画像用于游戏难度自适应并开展人类研究验证。
AI 中文摘要
个性化游戏生成需要从玩家的游玩方式中推断其能力与行为风格。大型语言模型(LLM)使这种推断比以往更易实现:LLM可读取原始游戏游玩记录,生成流畅、合理的玩家画像。然而,合理性未经验证,而验证正是该领域所缺失的:潜在特质不可观测;问卷提供的是带有噪声的替代指标,且当用自我报告验证基于行为的推断时会陷入循环;若缺乏情境,行为本身也存在歧义——从不收集道具的玩家可能是不需要,也可能是从未有机会获得。我们解决了这两个问题。首先,我们构建了一个合成玩家群体,其特质本质上是真实的:每个特质都是显式的bot参数,仅在受控操作产生一致的、特质特定的行为变化后才被接受。与先前反转已知决策模型的参数恢复工作不同,我们的基准仅从行为记录评估与策略无关的推断。其次,我们引入了机会感知的决策时刻表示,将偏好与表达偏好的机会解耦;选择性移除该表示会降低依赖机会的特质的表现。在该基准上,少样本LLM推断在大多数特质上优于基于嵌入和规则的基线,尽管基于特征的监督回归模型总体上仍更强。最后,我们闭环验证:推断出的画像驱动难度自适应,与真实参考及不匹配画像对照组进行评估,且一项探索性人类研究检验这些发现是否可迁移至真实玩家。
英文摘要
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable; questionnaires provide noisy proxies and become circular when self-reports are used to validate behavior-based inference; and behavior itself is ambiguous without context -- a player who never collects an item may not want it, or may never have had the chance. We address both problems. First, we construct a synthetic player population whose traits are ground truth by construction: each trait is an explicit bot parameter, accepted only after controlled manipulation produces consistent, trait-specific behavioral change. Unlike prior parameter-recovery work that inverts a known decision model, our benchmark evaluates policy-agnostic inference from behavioral transcripts alone. Second, we introduce an opportunity-aware decision-moment representation that disentangles preference from the chance to express it; ablating it selectively degrades opportunity-dependent traits. On this benchmark, few-shot LLM inference outperforms embedding- and rule-based baselines on most traits, though feature-based supervised regressors remain stronger overall. Finally, we close the loop: inferred profiles drive difficulty adaptation, evaluated against ground-truth references and mismatched-profile controls, and an exploratory human study examines whether these findings transfer to real players.
Comments16 pages, 3 figures, 6 tables. Includes technical appendix