arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于重分配的代价推断改进稀疏安全离线强化学习

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Ebenezer Gelo, Geraud Nangue Tasse, Steven James, Benjamin Rosman

arXiv 2608.12306首次发表:更新:

发表机构

University of the Witwatersrand(威特沃特斯兰德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对安全离线强化学习中稀疏轨迹级停止反馈问题,提出RCI框架转换为密集每步代价,经理论证明转换无损,实验在两类任务上违规率更低且鲁棒性好。

AI 中文摘要

安全离线强化学习通常假设可获取密集的每步代价标注,但实际中监督者仅提供轨迹级的停止反馈:在首次不安全转换处的二元信号,无每步归因。我们将此视为时间信用分配问题,提出基于重分配的代价推断(Redistribution-based Cost Inference,RCI)框架,通过回报分解将稀疏停止反馈转换为密集的每步代价,随后在增强数据集上训练受约束的离线策略。我们证明,回报等价重分配在约束马尔可夫决策过程(CMDP)中保留可行策略集和最优拉格朗日量,从理论上确立该转换是无损的,同时在实践中改善代价评论者学习的条件。在高速公路驾驶和机器人操纵任务上的实验表明,与稀疏方法和基于分类器的基线相比,其违规率显著更低,且对异构数据集组成和标签噪声具有鲁棒性。

英文摘要

Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset. We show that return-equivalent redistribution preserves the feasible policy set and the optimal Lagrangian in a CMDP, establishing that the transformation is lossless in theory while yielding better-conditioned cost critic learning in practice. Experiments on highway driving and robotic manipulation demonstrate substantially lower violation rates than sparse and classifier-based baselines, with robustness to heterogeneous dataset compositions and label noise.

CommentsAccepted at the 1st IJCAI Workshop on Safe Physical AI (SPAI 2026), affiliated with IJCAI/ECAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑