arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07719cs.LG

CODS:用于可复用离线强化学习的迭代贝尔曼残差数据选择

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对离线强化学习中冗余数据成本高、简单子采样易丢失关键转移的问题,提出可复用的CODS数据选择方法,在D4RL、ALFWorld、GSM8K等任务上取得优异性能。

中文摘要 AI 辅助

离线强化学习需从固定转移池中反复训练策略,冗余数据会在不同随机种子和超参数设置下增加计算成本,而简单的子采样可能会移除长视野信用分配所需的稀有转移。本文提出CODS,一种由评论者(critic)引导的选择器,其交替进行与算法匹配的评论者拟合和高残差转移的获取,随后冻结一个可复用子集。与优先回放(prioritized replay)不同,CODS生成静态人工产物;与一次性残差选择不同,它会随评论者变化更新分数。在10%的预算下,CODS在20个有效的D4RL任务-算法单元中保留了合格池性能的96.6%,在19/20个单元上优于ReDOR和OPER,在20/20个单元上优于所有其他子集基线;所有6种子集优势在采用Holm校正的预先声明的分层推理下仍显著。在保持选择器总更新次数固定的情况下,5次获取轮次比1次获取轮次在4个代表性单元上提升了11.23个点,之后趋于饱和。等迭代次数和等时间评估表明,是复用而非单次运行的加速带来了计算优势。机制和干扰实验揭示了有用的稀疏奖励增强以及对异常值的敏感性。最后,全轨迹扩展保留了池化ALFWorld成功率的95.4%和池化GSM8K精确匹配率的96.5%。因此,CODS是一种可复用的选择程序,而非正式的核心集(coreset)保证。

英文摘要

Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment. We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset. Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes. At a 10\% budget, CODS retains 96.6\% of eligible-pool performance across 20 valid D4RL task--algorithm cells. It exceeds ReDOR and OPER on 19/20 cells and every other subset baseline on 20/20; all six subset advantages remain significant under predeclared hierarchical inference with Holm correction. Holding total selector updates fixed, five acquisition rounds improve four representative cells by 11.23 points over one round and saturate thereafter. Equal-pass and equal-hour evaluations clarify that reuse, rather than a single-run speedup, creates the compute advantage. Mechanism and corruption interventions expose both useful sparse-reward enrichment and sensitivity to outliers. Finally, a whole-trace extension retains 95.4\% of pooled ALFWorld success and 96.5\% of pooled GSM8K exact match. CODS is therefore a reusable selection procedure, not a formal coreset guarantee.

发表机构

  • Iowa State University(爱荷华州立大学)
  • BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

↑