arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性函数逼近下带分段奖励反馈的强化学习

Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation

Fengxu Liu, Siwei Wang, Gal Dalal, Shie Mannor, Yihan Du

arXiv 2610.08271首次发表:更新:

发表机构

National University of Singapore; Microsoft Research Asia; NVIDIA Research; Technion; Singapore University of Technology and Design(新加坡国立大学; 微软亚洲研究院; 英伟达研究院; 以色列理工学院; 新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究线性函数逼近下分段奖励反馈的强化学习,提出针对二元和求和反馈的算法,发现二元反馈下分段数指数降低遗憾,求和反馈下粒度影响小,等长分段最优。

AI 中文摘要

经典强化学习(RL)假设每个访问的状态-动作对都能观察到奖励。然而,在诸如自动驾驶等实际应用中,这种细粒度的反馈可能代价高昂或难以收集,而轨迹级别的反馈对于高效学习而言可能过于稀疏。为了提供一种介于这两者之间的通用反馈模型并处理大规模状态空间,我们研究了在线性函数逼近下带有分段奖励反馈的强化学习。我们的工作回答了分段反馈的粒度以及分段的选择如何影响学习。对于具有已知转移概率的等长分段,我们分别针对二元反馈和求和反馈类型设计了算法\bitssegd\\和\edlinucbsegd\\。它们采用带规划的后验采样以实现计算效率,并采用E-最优实验设计以达到接近最优的性能。我们建立了几乎匹配的下界。对于具有未知转移概率的等长分段,我们开发了一个统一的\seglsvits\\框架,并针对二元反馈和求和反馈提供了两种实例化,该框架将后验估计的奖励参数仔细地整合到最小二乘值迭代中。这些结果揭示了一个基本见解:在二元反馈下,增加分段数量通过一个指数因子显著减少了遗憾值,而令人惊讶的是,在求和反馈下,分段的粒度对学习影响不大。最后,为了研究根据状态-动作特征进行分段是否能进一步加速学习,我们设计了一种允许任意分段的算法\uneqsegbitsd\\。由此产生的遗憾界表明,在通常的椭圆势分析下,状态-动作特征对遗憾值的影响仅通过对数因子体现,并且等长分段实现了最佳性能。

英文摘要

Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning. To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation. Our work answers how the granularity of segment feedback and the choice of segmentation influence learning. For equal-length segments with known transitions, we design algorithms $\bitssegd$ and $\edlinucbsegd$ for binary and sum feedback types, respectively. They adopt posterior sampling with planning to achieve computational efficiency and the E-optimal experimental design to attain near-optimality. Nearly matching lower bounds are established. For equal-length segments with unknown transitions, we develop a unified $\seglsvits$ framework with two instantiations for binary and sum feedback, which carefully integrates the posterior estimated reward parameters into least-squares value iteration. These results reveal a fundamental insight: under binary feedback, increasing the number of segments significantly reduces the regret through an exponential factor, while surprisingly, under sum feedback, the granularity of segments does not affect learning much. Finally, to investigate whether segmenting according to state-action features can further expedite learning, we design an algorithm $\uneqsegbitsd$ that allows arbitrary segmentations. The resulting regret bound shows that under the usual elliptical potential analysis, the influence of state-action features on the regret appears only through logarithmic factors, and equal segmentation achieves the best performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑