认证安全策展:安全离线强化学习的无分布保证
Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对仅能通过片段比较和偶发预算询问判断安全性的离线强化学习,提出认证安全策展流程,以无分布保证选择安全轨迹并克隆策略,在DSRL任务上多数满足预算且可预测拒绝。
中文摘要 AI 辅助
安全离线强化学习假设每个转移都带有代价函数。我们探究当安全只能通过比较短片段并偶尔询问某个回合是否超出预算来判断时,仍能实现什么。认证安全策展通过一个“先过滤后克隆”的流程给出答案:一个仅基于状态的、从片段比较中训练的价值函数对整条轨迹进行评分,Learn-then-Test校准在无分布的$(\alpha, \delta)$界下对选择的不安全比例进行认证,从而确定选择阈值,随后进行行为克隆。我们未发现先前工作对离线强化学习或模仿学习的训练集组成进行认证。Oracle控制验证了设计合理性:即使有精确的价值函数,对单个转移进行重新加权也会失败,因此价值函数选择整条轨迹。策略在十五个DSRL任务中的十一个上满足代价预算,比克隆真实安全子集(该子集需要对每条轨迹进行标注)少一个;未认证的变体达到十二个。在认证选择上重新训练后,最强的全标签方法变得安全,而其自身代价目标无论怎么设置都无法挽救。拒绝是可预测的:证书的概率在池达到的纯度上具有封闭形式,校准样本估计该纯度,而评分器仅通过该纯度进入。
英文摘要
Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(α, δ)$ bound on the unsafe fraction of the selection, and behavior cloning follows. What is certified is the training set, not the policy, whose cost we report rather than bound. Where no threshold attains the target, the procedure refuses. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on twelve of fifteen DSRL tasks, matching a clone of the ground-truth safe subset, which needs a label on every trajectory. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity of the pool's top quantile, which the calibration sample estimates and through which the scorer enters.
发表机构
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。