发表机构
North Carolina State University(北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在奖励和状态受对抗性损坏反馈下学习最优策略问题,提出基于批处理的BR-Async-Q算法,通过划分数据流为批次及构建鲁棒估计来抵御数据损坏,给出高概率误差界,为异步Q学习提供首个鲁棒性保证。
AI 中文摘要
受恶劣环境中强化学习的启发,我们考虑在对抗性损坏反馈下学习最优策略的问题。具体而言,在每个时间步,对手可根据Huber污染模型干扰学习者的奖励和状态观察。为抵御此类数据损坏,我们提出BR-Async-Q:一种基于两个关键思想的新颖的、基于轮次的鲁棒Q学习算法:(i)将在线数据流划分为批次以减少方差,(ii)使用此类批处理数据构建Bellman最优性算子的鲁棒估计。我们证明了BR-Async-Q的高概率\(\ell_\infty\)误差界,与普通Q学习的误差界相匹配,最多有一个与损坏样本比例成比例的小附加项。据我们所知,这为受奖励和状态损坏影响的异步Q学习提供了首个鲁棒性保证。此外,当只有奖励被损坏时,我们算法的界对损坏比例的依赖性是极小极大最优的。
英文摘要
Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observations of the learner following the Huber contamination model. To defend against such data corruption, we propose BR-Async-Q: a novel, epoch-based, robust Q-learning algorithm built upon two key ideas: (i) partitioning the online data stream into batches to reduce variance, and (ii) constructing robust estimates of the Bellman optimality operator using such batched data. We prove a high-probability $\ell_\infty$ error bound for BR-Async-Q that matches that for vanilla Q-learning, up to a small additive term that scales with the fraction of corrupted samples. To our knowledge, this provides the first robustness guarantee for asynchronous Q-learning subject to both reward and state corruption. Furthermore, when only rewards are corrupted, the dependence of our algorithm's bound on the corruption fraction is minimax optimal.
CommentsTo appear at the 65th IEEE Conference on Decision and Control (CDC) 2026