arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40221cs.LGcs.AIcs.CL

PhantomEnvironments:在虚构世界中训练LLM智能体

PhantomEnvironments: Training LLM Agents in Fictional Worlds

Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q. Weinberger

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PhantomEnvironments,利用纯规则生成的虚构世界环境训练LLM搜索智能体,无需LLM参与生成且零成本,训练出的智能体可迁移至真实多跳搜索基准,并展现出随问题难度线性扩展的搜索预算能力。

中文摘要 AI 辅助

使用强化学习(RL)训练LLM智能体受到环境的瓶颈制约,环境必须提供可验证的奖励、支持长时程交互,并且能够廉价地扩展。现有方法依赖于昂贵的人工策划数据,或依赖于LLM生成的环境,这些环境存在幻觉和基准污染的风险。我们证明,LLM可以通过完全由规则生成的合成环境被训练成有能力的搜索智能体,这种环境的生成不需要LLM,且边际成本为零。我们构建了PhantomEnvironments,即来自虚构世界的多轮RL环境,其中智能体必须搜索模板化文章语料库来回答多跳问题。尽管这些环境与现实世界不共享任何事实,但这些极其简单的环境却能产生可迁移到现实世界多跳搜索基准的智能体,在较新的基准上往往优于使用现实世界训练数据训练的智能体。经过训练的智能体能够泛化到未见过的虚构宇宙,并且Qwen模型学会将其搜索预算大致线性地随问题难度扩展,这表明仅通过环境交互就能涌现出搜索扩展能力。对环境复杂性的消融研究表明,跳数对迁移的驱动作用大于约束或比较:即使是最简单的规则生成环境,也是训练可泛化LLM智能体的一个出奇有效且免费的资源。

英文摘要

Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.

发表机构

  • Cornell University(康奈尔大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑