arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EVOKE:在智能体中激发世界知识以实现可迁移决策

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng

arXiv 2609.38334首次发表:更新:

发表机构

Shanghai Jiaotong University; Eastern Institute of Technology; Zhongguancun Academy; Hong Kong Polytechnic University(上海交通大学; 东方理工高等研究院; 中关村学院; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EVOKE通过固定状态下目标多样性激发LLM智能体内化的世界知识,提升未见环境中的决策迁移能力,实验显示任务性能、泛化和数据效率均提升。

AI 中文摘要

大型语言模型(LLMs)越来越多地被部署为智能体,用于多步决策,但在未见过的环境中迁移能力较差。世界模型方法通过训练智能体预测未来观测来解决这一问题,但代价是额外的训练以及预测用于规划时误差的累积。然而,对于在数字环境中运行的LLM智能体,大部分世界知识在预训练期间已被内化,这便将问题从获取知识转变为激发知识。我们认为,典型的后训练对这类激发提供的压力很小,因为在每个访问状态下的单一目标监督无意中驱使策略依赖肤浅的上下文习惯。我们引入EVOKE,一种后训练方法,通过在固定状态下提供目标多样性来施加这种压力。受理论启发——该理论表明,一个在多样目标下胜任的智能体必须编码一个可从其动作偏好中恢复的世界模型——EVOKE保持环境状态和交互历史固定,并在替代目标下对相同的候选动作进行排序,迫使动作偏好发生变化,从而使依赖上下文习惯或单一目标相关性的策略无法正确排序。这隐式地激发策略预训练的世界知识来指导决策。我们在三个骨干网络中跨多样任务评估EVOKE,展示了改进的任务性能、未见环境泛化和数据效率。我们进一步进行受控分析,以更好地理解这些收益的驱动因素。这些发现为通过直接决策监督激发内化世界知识以实现可迁移动作提供了新视角。

英文摘要

Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.

Comments19 pages. Project page: https://gnonymous.github.io/EVOKE ; Code: https://github.com/Gnonymous/EVOKE ; Models: https://huggingface.co/Gnonymous/EVOKE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑