世界模型与策略如何在大语言模型智能体中结合?一种联合谱分析与行为学解释
How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account
浏览论文内容
中文总结 AI 辅助
该研究通过受控实验探究LLM智能体中世界模型与策略的结合机制,从几何与行为角度揭示二者的互补特性,并提出提升策略训练保留世界知识能力的方法。
中文摘要 AI 辅助
大语言模型(LLM)智能体如何既理解其所处的环境,又掌握其中设定的任务?我们结合世界模型训练(下一状态预测)与策略训练(奖励最大化)开展受控实验,探究该问题。我们通过加性参数更新剖析所得模型,从几何角度发现:无论分开训练还是顺序训练,有效的世界模型更新均为低秩,且与策略更新共享输入特征子空间,但写入近乎正交的输出方向。然而,投影干预实验显示,在移除世界模型的主导输入方向时,顺序更新的策略强化学习(RL)比分开训练的策略更具鲁棒性,表明其已习得替代输入通路。从行为角度发现,顺序训练的智能体探索更广泛的状态与动作范围。基于此,我们提出问题:策略训练能否尽可能保留世界知识?我们采用基于几何驱动输入基的无训练合并方法,结合策略RL期间的在线世界模型损失进行探究,结果显示两种方法均优于未处理的基线。我们的发现表明,世界知识与任务导向能力可按几何互补形式习得,未来的后训练流程应考虑如何最佳构建二者间的接口。
英文摘要
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
发表机构
- Dartmouth College(达特茅斯学院)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。