先让它可玩,再让它变好:面向小型对话游戏智能体的分阶段交互学习
First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents
查看机构详情
- City, University of London(伦敦城市大学)
- The Alan Turing Institute(艾伦·图灵研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出Qwen-GuidePlay-2B模型,通过分阶段微调方法训练小型对话游戏智能体,在Playpen挑战赛中取得优异成绩,证明精心筛选策略可让小型模型实现高性能。
中文摘要 AI 辅助
我们提出Qwen-GuidePlay-2B,这是一个用于对话游戏交互的2B参数语言模型。我们分三个步骤对Qwen3.5-2B进行微调:a)仅在Playpen中成功的游戏轨迹上进行监督微调(SFT);b)加权回合级监督微调;c)教师引导的监督微调。教师模型是一个更大的模型,仅用于修正格式和评估示例,不生成新的黄金动作。我们的最终模型在公开的Playpen验证集上取得了57.12的clemscore和42.68的statscore。在官方发布的挑战赛结果中,我们的模型在提交系统中获得了第二高的Playpen clemscore增量,约比其基础模型高+36。我们的发现表明,模仿完整轨迹有助于提升可玩度,而回合级和教师引导的训练通常能改善决策并提高整体得分。像重播修复和难例挖掘这类程序上繁重的替代方法并无帮助,这表明小型模型仅需通过精心的筛选策略而非激进的改动即可实现高性能。我们公开了模型和代码以支持可复现性。
英文摘要
We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guided SFT. The teacher model (which is a larger model) is only used to fix formatting and evaluate examples, but does not create new gold actions. Our final model scores 57.12 clemscore and 42.68 statscore on the public Playpen validation. In the officially released challenge results, our model obtains the second-highest Playpen clemscore delta among submitted systems (which is approximately +36 over its base model). Our findings suggest that imitating full trajectories helps with playability, while turn-level and teacher-guided training usually improve decision-making and increase the overall score. Alternative procedurally heavy approaches like replay-repair and hard-example mining did not help, which suggests that small models are performant simply by using careful curation strategies rather than aggressive changes. We make available both the model and the code for reproducibility.