arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

设计安全:带连续动作的上下文博弈机的实现代价约束

Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions

Spyros Dragazis, Aldo Pacchiano

arXiv 2608.26755首次发表:更新:

发表机构

Boston University; Broad Institute of MIT and Harvard(波士顿大学; 麻省理工学院与哈佛大学博德研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对带连续动作的上下文博弈机,提出高概率约束UCB算法,通过保守估计安全动作集保障实现代价安全,实验显示其能显著减少安全违规。

AI 中文摘要

上下文博弈机是不确定性下序贯决策的标准框架,应用于临床试验、剂量选择、推荐系统和自主系统等领域。安全是这些应用的核心,因为剂量选择或自动驾驶等场景中单次不安全决策可能造成灾难性后果。博弈机问题中建模安全的常用方式是为每个动作关联奖励信号和代价信号,在代价约束下优化奖励。现有多数安全约束博弈机模型通过要求每个动作的期望代价低于规定阈值来保障安全,但在异方差场景中这可能不足——所选动作不仅影响期望奖励和代价,还影响观测结果的变异性。本文研究带一维连续动作、阶段式实现代价高概率约束的上下文博弈机,提出高概率约束UCB(High-Probability Constrained UCB),这是一种乐观-悲观算法,在探索奖励的同时保守估计安全动作集。针对线性奖励和代价模型,本文证明紧的$\tilde{\text{O}}(d"T)$遗憾界,并利用Eluder维度将分析扩展到一般函数类。实验表明,与期望代价约束基线相比,实施实现代价安全能大幅减少违规情况。

英文摘要

Contextual bandits are a standard framework for sequential decision-making under uncertainty, with applications in clinical trials, dosage selection, recommendation systems, and autonomous systems. Safety is central in many of these applications, since a single unsafe decision in settings such as dosage selection or autonomous driving can have catastrophic consequences. A common way to model safety in bandit problems is to associate each action with both a reward signal and a cost signal, and to optimize reward subject to constraints on cost. Most existing safety-constrained bandit models enforce safety by requiring the expected cost of each action to remain below a prescribed threshold. However, this may be insufficient in heteroscedastic settings, where the chosen action affects not only the expected reward and cost, but also the variability of the observed outcomes. We study contextual bandits with one-dimensional continuous actions and stage-wise high-probability constraints on the realized cost. We propose High-Probability Constrained UCB, an optimistic-pessimistic algorithm that explores for reward while conservatively estimating the safe action set. For linear reward and cost models, we prove a tight $\tilde{\mathcal{O}}(d\sqrt{T})$ regret bound, and we extend the analysis to general function classes using the eluder dimension. Experiments show that enforcing realized-cost safety substantially reduces violations compared with expected-cost constrained baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑