可验证的隐藏动态玩法:从已求解机制生成智能体强化学习环境
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
查看机构详情
- Georgia Institute of Technology(佐治亚理工学院)
- Alibaba Token Foundry, Alibaba Group(阿里巴巴集团阿里通义实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
VHD-Play 通过先求解数学模型再生成有状态环境,以低成本产出多样智能体训练环境,显著提升智能体性能并泛化至外部基准。
中文摘要 AI 辅助
语言模型智能体日益面临具有演化状态、相互依赖决策和延迟结果的长时程任务。扩展其训练需要多样化的智能体环境、可靠的结果信号和较低的扩展成本。现有的生成流程通常先构建环境,再定义其结果规则或标注其轨迹,导致动态过程和评估事后才对齐。VHD-Play 颠倒了这一依赖关系:先采样并求解一个数学模型,再由基于语料库的设置器将其决策过程渲染为有状态工具。可执行的动态过程和轨迹评分参考均继承自同一已求解模型。该流程以每个环境几美分的成本生成了 3,300 个多样化的智能体环境。在三个环境族上训练 Qwen3.6-35B-A3B,将其在五族诊断中的平均智能体得分从 0.204 提升至 0.815。在来自全部三个训练族和八个未见机制族的保留实例上也出现了提升,并进一步扩展到生成基质之外的外部基准,涵盖通用函数调用、旅行规划和 365 天电子商务。在 E-Commerce Bench 上,训练后的检查点在不破产的情况下完成了每次运行,并超过了 Qwen3.7-Max。我们比较了书面描述的问题与揭示或隐藏其参数的有状态版本。比较表明,大部分可学习的差距在于有状态交互,而非底层问题求解。冻结的 35B 设置器可实现更大的环境,且规模匹配的训练在机制规模和时域增长时保持收益,表明演化训练基质的潜力。
英文摘要
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.