从预测到决策:面向连续铲装作业的世界模型引导动作选择
From Prediction to Decision: World-Model-Guided Action Selection for Continuous Pile Excavation
浏览论文内容
中文总结 AI 辅助
针对轮式装载机连续铲装问题,提出世界-动作模型(WAM),通过预测地形变化与装载量并排序候选铲取动作,将平均铲取次数降低17.1%,并实现全尺寸机器上的闭环自主作业。
中文摘要 AI 辅助
轮式装载机铲装是一个序贯决策问题,其中每一次铲取都会改变后续动作可用的地形。一个实用的世界模型必须准确预测动作后果,实时对候选动作进行排序,并在全尺寸机器的闭环内运行。我们提出了世界-动作模型(WAM),该模型提出多个铲取候选,拒绝几何上不可行的候选,联合预测有符号的地形变化和装载体积,执行预测装载量最大的候选,并根据新观测到的地形重新规划。在32个几何不相交的MinSlope测试场景中,将世界模型排序添加到匹配的扩散候选生成中,将平均铲取次数从651.8次减少到540.6次(减少17.1%),保持了32/32的完成率,并改善了每一对场景的表现。在完整系统比较中,WAM完成了32/32个场景,而独立训练的软演员-评论家策略仅完成29/32个场景。对输入表示、空间支持和五种架构的比较,确定了一种准确且高效的物理结构预测器。我们进一步在事件不相交的全尺寸装载机数据上评估了该接口,并部署了完整的感知-提议-预测-选择-执行循环用于自主铲装。ROS2/TensorRT实现可在Jetson AGX Orin上以72.4毫秒处理五个候选。仿真结果确立了决策层面的增益,而物理实验则证明了现实世界闭环的可行性。
英文摘要
Wheel-loader excavation is a sequential decision problem in which every scoop changes the terrain available to subsequent actions. A practical world model must predict action consequences accurately, rank candidates in real time, and operate inside the closed loop of a full-size machine. We present the World-Action Model (WAM), which proposes multiple scoops, rejects geometrically inadmissible candidates, jointly predicts signed terrain change and loaded volume, executes the candidate with the largest predicted load, and replans from the newly observed terrain. On 32 geometry-disjoint MinSlope test episodes, adding world-model ranking to matched diffusion proposals reduces the mean scoop count from 651.8 to 540.6 (17.1%), preserves 32/32 completion, and improves every paired episode. In a complete-system comparison, WAM completes 32/32 episodes versus 29/32 for an independently trained soft actor-critic policy. Comparisons of input representations, spatial support, and five architectures identify an accurate and efficient physics-structured predictor. We further evaluate the interface on event-disjoint full-size-loader data and deploy the complete perception-proposal-prediction-selection-execution loop for autonomous excavation. The ROS2/TensorRT implementation processes five candidates in 72.4 ms on a Jetson AGX Orin. The simulation results establish decision-level gains, while the physical experiments demonstrate real-world closed-loop feasibility.
发表机构
- Tsing-AI (Shanghai) Technology Co., Ltd.(清智(上海)科技有限公司)
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Ocean University(上海海洋大学)
机构由 AI 辅助整理,请以论文原文为准。