arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeliveryGym:一种用于具有自适应课程的长时程具身智能体规划的可执行强化学习环境

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

Haoqiang Kang, Yiming Zhang, Yiyang Guo, Chuying Li, Jianzhi Shen, Tianruo Rose Xu, Xiaokang Ye, Lianhui Qin

arXiv 2609.19801首次发表:更新:

发表机构

UC San Diego; Johns Hopkins University; Cornell University(加州大学圣迭戈分校; 约翰斯·霍普金斯大学; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeliveryGym是一个3D强化学习环境,通过模拟连续快递轮班中的资源依赖和轨迹奖励,训练具身智能体进行长时程规划,实验表明RL提升净收入54.3%,自适应课程进一步改善测试性能。

AI 中文摘要

可执行环境使大语言模型(LLM)智能体能够从自身行为的后果中学习。对于具身智能体而言,这些后果不仅限于当前任务是否成功:完成一次配送可能会消耗后续工作所需的时间、能量或金钱。因此,学习规划需要能够保留这些依赖关系并将其转化为整个轨迹中反馈的环境。我们引入了DeliveryGym,这是一个用于在连续快递员轮班中评估和训练智能体的3D环境。它将多模态工具交互与持久的世界动态相结合,并从模拟器事件计算轨迹奖励,从而使智能体决策的成本可用于强化学习(RL)。该环境还会根据策略观察到的弱点调整未来的训练轮班,同时保持评估固定。在六个模型和13个城市地图上,评估揭示了可靠执行分配配送与在轮班中选择和排序工作之间的差距。在固定测试套件上,RL将Qwen3-VL-4B的净收入提高了54.3%,表明在耦合约束下从完整轮班中学习能提升性能。在相同回滚预算下,调整训练环境比均匀采样使测试收入提高了16.5%,表明智能体练习的情境也很重要。DeliveryGym提供了一个可执行环境,用于研究智能体如何在一个情节内学习协调配送并为后续订单保留资源。

英文摘要

Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, or money needed for later work. Learning to plan therefore requires environments that preserve these dependencies and turn them into feedback across a complete trajectory. We introduce DeliveryGym, a 3D environment for evaluating and training agents on continuous courier shifts. It couples multimodal tool interaction with persistent world dynamics and computes trajectory rewards from simulator events, making the costs of an agent's decisions available for reinforcement learning (RL). The environment also adapts future training shifts to the policy's observed weaknesses while keeping evaluation fixed. Across six models and 13 city maps, evaluation exposes a gap between reliably executing assigned deliveries and choosing and sequencing work over a shift. On the unseen-city test set, RL improves Qwen3-VL-4B's net income by 54.3%, showing that learning from complete shifts improves performance under these coupled constraints. Adapting the training environment improves evaluation income by 18% over uniform sampling at 100 updates, indicating that which situations an agent practices also matters. DeliveryGym provides an executable setting for studying how agents learn to coordinate deliveries and preserve resources for later orders within an episode.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑