发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出交错投影梯度下降的模仿学习方法,在约束下训练神经网络控制器,通过交替安全投影步骤实现约束满足,在赛车任务中降低违反率至1%且不牺牲性能。
AI 中文摘要
我们提出了一种在状态和输入约束下用于神经网络控制策略的模仿学习设计。训练过程交替进行标准模仿梯度步骤与包含$k$个安全步骤的块,这些安全步骤将网络的行动拉向其到安全集的投影;在运行时,控制器仅为训练后的网络,无需安全过滤器。我们将该方案分析为策略行动空间中的非精确投影梯度下降。当投影行动在每个安全步骤中重新计算且每个步骤一致地将行动移向安全集时,让$k$以对数方式增长可在训练状态上实现渐近约束满足,并将与模仿损失约束最优解的距离限定;若投影行动保持固定,则仅当它们能被网络精确表示时,上述结论才成立。在非线性自主赛车任务中,我们将我们的方法与向模仿损失添加加权约束违反惩罚的方法进行比较。在足够大的权重下,我们的方法匹配无约束模仿的单圈时间,同时将违反情节的比例从$15\%$降至$1\%$,约为惩罚方法在其最佳权重下的六分之一。其单圈时间对权重不太敏感,而权重反而设定训练期间违反消失的速度。在赛车中,安全修正稀疏且分析条件不成立;收益反而来自训练期间收集的数据。这些收益以额外训练计算为代价。
英文摘要
We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of $k$ safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting $k$ grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from $15\%$ to $1\%$, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
CommentsSubmitted to the 2027 American Control Conference (ACC). 9 pages, 3 figures