发表机构
Patronus AI(守护神人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究强化学习中世界模型问题,对比自回归语言模型和掩码扩散语言模型,发现掩码扩散语言模型性能更优。引入GRPO训练框架,在三个OOD环境上零样本转移消融效果良好,还进行相关分析与评估,最后开源工作推动该领域研究。
AI 中文摘要
强化学习的发展需要多样化、专门的训练环境。手工制作的固定任务和奖励环境随着模型性能提升而失效,长期稀疏奖励会导致模式崩溃。世界模型可模拟环境状态,有潜力扩展多样性,但自回归世界模型存在从左到右的偏差。本文将基于文本的世界建模形式化为可控的过渡动力学问题,整理了跨越九个开源环境和十二个前沿模型家族的239,403个基于现实的状态-动作轨迹。比较了自回归语言模型和掩码扩散语言模型,表明掩码扩散语言模型通过双向锚点感知去噪,在可比推理延迟下比参数规模大四倍的语言模型具有更好的连贯性、基于现实性和经验验证的展开多样性。引入了具有确定性状态检查的即插即用GRPO训练框架,在三个OOD环境上进行零样本转移消融,在无特定环境微调的情况下比基线实现高达47%的绝对增益。还对对抗场景下的失败模式进行了行为分析,并进行了人类对现实性、结果正确性和训练效用的评估。最后开源工作以鼓励该方向的研究。
英文摘要
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial state, task context, tool schemas, domain rules, and steering directives, and (ii) curate 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families. We compare AR LMs and masked diffusion language models (MDLMs), showing MDLMs, via bidirectional anchor-aware denoising, achieve better coherence, groundedness, and empirically validated rollout diversity than LLMs over 4x their parameter size, at comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three OOD environments (ScienceWorld, ALFWorld, AppWorld) across three 1.2B-7B agent backbones (LFM2.5, Qwen3, Mistral), achieving up to 47% absolute gains over baselines without environment-specific fine-tuning. We further conduct behavioral analysis of failure modes under adversarial scenarios and human evaluation on realism, outcome correctness, and training utility. We open-source our work to encourage research in this direction.
CommentsAccepted to NeurIPS 2026 Dataset: https://huggingface.co/PatronusAI/world_model_corpus Training code: https://github.com/patronus-ai/mdlm_world_modeling