arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20098cs.LG

CARE-VI:离策略演员-评论家学习中用于价值提升的保守自适应可靠性估计

CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo

AI总结:

针对离策略演员-评论家学习中价值提升目标不可靠的问题,提出CARE-VI框架,通过CARS、SEVA和DARE三个组件实现证据调节的目标构建,在四个MuJoCo任务上取得最优平均回报。

AI中文摘要:

可靠的时间差分目标是离策略演员-评论家学习的核心。直接价值提升通过备选动作细化下一状态目标,但这种细化的可靠性取决于候选动作如何被排序、审查和加权。噪声排序可能迫使过早的候选承诺,重用选择分数可能使目标估值产生偏差,而固定的增强权重可能放大弱证据。为应对这些风险,我们开发了保守自适应排序与筛选(CARS),它在预设预算内保留有序候选前缀,并且仅当观测到的边界间隙超过按分歧缩放的置信半径时才缩小该前缀。选择器-评估器价值评估(SEVA)使用选择器评论家对候选进行排序,并使用独立参数化的评估器评论家审查所选价值,然后将审查后的价值限制在选择器参考值之内。动态自适应风险感知增强(DARE)随后利用候选可靠性、选择器与评估器信号之间的间隙以及有限阶段因子来调节每个残差修正。CARS、SEVA和DARE共同构成CARE-VI,一个证据调节的目标构建框架,该框架保留了用于评论家回归和演员更新的骨干接口。分析限定了CARS边界误差、SEVA所选价值高估以及DARE残差位移与其总体对应物的单侧偏差,并确立了有限阶段扰动结束后固定策略的恢复。在四个MuJoCo任务上使用SAC、TD3和TD7进行的实验表明,CARE-VI在所有十二种设置中均取得了最高平均回报。分组消融和标量诊断支持了这三个组件在提高目标可靠性方面的作用。

英文摘要:

Reliable temporal-difference targets are central to off-policy actor-critic learning. Direct value improvement refines the next-state target with alternative actions, but the reliability of this refinement depends on how candidate actions are ranked, reviewed, and weighted. Noisy rankings may force premature candidate commitment, reusing selection scores may bias target valuation, and fixed enhancement weights may amplify weak evidence. To address these risks, we develop Conservative Adaptive Ranking and Screening (CARS), which retains an ordered candidate prefix within a preset budget and narrows it only when the observed boundary gap exceeds a disagreement-scaled uncertainty radius. Selector-Evaluator Value Assessment (SEVA) uses selector critics to order candidates and a separately parameterized evaluator critic to review the selected value, then caps the reviewed value at the selector reference. Dynamic Adaptive Risk-aware Enhancement (DARE) then regulates each residual correction using candidate reliability, the gap between selector and evaluator signals, and a finite stage factor. Together, CARS, SEVA, and DARE form CARE-VI, an evidence-regulated target construction framework that preserves the backbone interfaces for critic regression and actor updates. The analysis bounds the CARS boundary error, the SEVA selected-value overestimation, and the one-sided deviation of the DARE residual displacement from its population counterpart, and establishes fixed-policy recovery after the finite-stage perturbation ends. Experiments with SAC, TD3, and TD7 on four MuJoCo tasks show that CARE-VI achieves the highest mean return in all twelve settings. Grouped ablations and scalar diagnostics support the roles of the three components in improving target reliability.

↑