发表机构
National University of Singapore; Fudan University; Tsinghua University; The University of Sydney(新加坡国立大学; 复旦大学; 清华大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出世界潜能模型(WPM),利用预训练模型的世界知识评估长视界智能体的任务进展,无需任务特定微调,并在ALFWorld和ScienceWorld中显著提升策略优化成功率。
AI 中文摘要
长视界语言智能体通常仅从最终任务结果获得监督信号,这导致难以区分富有成效的中间行为与停滞甚至倒退。与其为每个任务单独学习价值函数或过程奖励模型,我们探究预训练模型能否利用其已有的世界知识来识别任务进展。我们将这一能力形式化为世界潜能模型(WPM),一种在智能体语境下对任务相对已实现进展进行目标条件评估的模型。在ALFWorld和ScienceWorld中,现成的预训练模型在无需针对特定任务进行评估器微调的情况下,在恢复已实现进展结构方面显著优于随机水平。我们进一步将这些进展判断锚定到任务特定的里程碑,以获得标量世界潜能,其时间差分可为策略优化提供过程敏感的步骤级信用。在匹配比较下,WPM引导的优化在所有评估配置中均优于仅基于结果的GRPO。综合这些结果,初步证据表明预训练世界知识能够支持可复用的已实现进展评估,并为长视界智能体提供有用的监督。
英文摘要
Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.