压缩驾驶策略中能力损失与恢复的闭环评估
A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
浏览论文内容
中文总结 AI 辅助
本研究提出分阶段闭环评估方法,在 Gym-Duckietown 中训练 PPO 策略后分阶段压缩驾驶策略,发现结构化剪枝致能力首次损失,蒸馏受数据限制,整数量化影响部分课程,未剪枝则保留全部课程,为自动驾驶安全部署提供实证分析。
中文摘要 AI 辅助
许多汽车与出行公司在内存和功耗有限的嵌入式计算机上部署学习得到的驾驶策略。剪枝、知识蒸馏和量化是减小这些策略规模并降低推理成本的标准方法。然而,这些方法通常通过 aggregate numerical scores(聚合数值分数)进行评估,而这类分数可能无法反映策略在与其他道路使用者交互时的安全驾驶能力。本研究提出一种分阶段闭环评估方法,以跟踪驾驶策略通过压缩流程的全过程。我们将驾驶任务建模为部分可观测马尔可夫决策过程(POMDP),并在 Gym-Duckietown 中使用近端策略优化(PPO)训练信念状态策略。随后,我们提取 actor 网络,分阶段对其进行压缩,并在五个驾驶课程上对其进行评估。研究表明,结构化剪枝是驾驶能力首次出现损失的阶段;与此同时,蒸馏可提升剪枝后的 actor 网络,但提升效果受其 rehearsal data(排练数据)的限制;对经蒸馏提升后的 actor 网络进行整数量化会导致车辆需要先停车再重启的部分课程失效;有趣的是,对未剪枝的 actor 网络执行相同流程则可保留全部五个课程。因此,本研究提供了一项实证分析,旨在回应当前关于如何判定压缩后的驾驶策略是否可接受的活跃讨论,以实现自动驾驶功能的安全且统计可靠的部署。
英文摘要
Many automobile and mobility companies deploy learned driving policies on embedded computers with limited memory and power. Pruning, knowledge distillation, and quantization are the standard methods to reduce the size and the inference cost of these policies. However, these methods are commonly assessed by aggregate numerical scores, and such scores may not reflect the ability of the policy to drive safely when interacting with other road users. In this study, we propose a stage-wise closed-loop evaluation approach to follow a driving policy through a compression pipeline. We formulate the driving task as a partially observable Markov decision process (POMDP) and train a belief-state policy with proximal policy optimization (PPO) in Gym-Duckietown. We then extract the actor, compress it one stage at a time, and evaluate it on five driving curricula. We show that structured pruning is the stage at which the driving capability is first lost. Meanwhile, distillation improves the pruned actor, but the improvement is limited by its rehearsal data. Integer quantization of the improved actor loses some of the curricula that require the vehicle to stop and then resume. Interestingly, the same procedure on the unpruned actor preserves all five curricula. Our study thus provides an empirical analysis aiming to answer the currently active discussions on how to accept a compressed driving policy, so as to achieve a safe and statistically reliable deployment of automated driving functions.
发表机构
- Universitas Muhammadiyah Yogyakarta(日惹穆罕马迪亚大学)
- Nara Institute of Science and Technology (NAIST)(奈良科学技术研究所)
- King Fahd University of Petroleum and Minerals (KFUPM)(法赫德国王石油与矿业大学)
机构由 AI 辅助整理,请以论文原文为准。