arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

获取、修复、保留:面向小模型对话游戏智能体的诊断引导式后训练方案

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

Nan Li

arXiv 2608.28458首次发表:更新:

发表机构

Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对小模型对话游戏智能体提出诊断引导式后训练方案,通过三步训练在LM Playschool Challenge中大幅提升了模型的域内对话游戏表现,同时保留了通用静态性能。

AI 中文摘要

交互式对话游戏测试了静态基准大多未明确体现的能力:模型必须跨轮次维护状态、解释反馈,并在不断变化的约束下选择有效动作。我们在LM Playschool Challenge中使用一个2B参数的开放权重模型研究该场景,发现许多失败不仅是广泛的知识失败,也是局部决策失败:重复猜测、格式错误的动作、以及违反模型刚收到的反馈。这些诊断结果催生了一个由三个步骤组成的训练方案:通过监督微调(SFT)获取广泛的游戏参与能力;使用针对单轮次的偏好对修复某一特定对话游戏类别内的可机械验证失败;保留超出这些对话游戏的通用能力。在官方最终评估中,我们的提交将公开Clemscore从10.67提升至38.92,封闭域内得分从13.41提升至41.17,同时大致保持了整体静态性能(基线为44.24,我们的结果为44.14)。域外Clemscore仍处于较低水平,为7.88,最大提升集中在目标类别的未见过变体中。我们的结果表明,广泛的SFT带来了模型的大部分能力提升;当失败检测精准时,单轮次监督可发挥有效作用,且观测到的迁移主要集中在类别内部。

英文摘要

Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of feedback that the model has just seen. These diagnostics motivate a training recipe organized around three steps: acquire broad game participation through supervised fine-tuning, repair mechanically verifiable failures within one targeted dialogue-game family using turn-local preference pairs, and preserve general capabilities beyond these dialogue games. In the official final evaluation, our submission improves public clemscore from 10.67 to 38.92 and closed in-domain score from 13.41 to 41.17, while approximately preserving aggregate static performance (44.14 vs. 44.24 for the baseline). Out-of-domain clemscore remains low at 7.88, with the largest gains concentrated in unseen variants of the targeted family. Our results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.

Comments14 pages, 14 tables; Accepted to the LM Playschool Workshop at EMNLP 2026; HF model card: https://huggingface.co/chnln/Qwen3.5-2B-playpen-playornotplay

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑