NFTR:从可证明的模式平均到离线目标条件强化学习中的测地线子目标选择
NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL
浏览论文内容
中文总结 AI 辅助
针对分层隐式Q学习在离线目标条件强化学习中存在的问题,提出NFTR方法,用条件归一化流取代高斯策略,结合基于架构三角不等式的三角松弛分数及RWDR目标,可避免高斯坍塌且在随机动力学下保持稳定。
中文摘要 AI 辅助
分层隐式Q学习(HIQL)是一种离线目标条件强化学习方法,仅通过价值函数优势来选择子目标。该规则有两种耦合的失败模式。乐观偏差将幸运的随机结果视为熟练的选择,而模式坍塌将多模态子目标分布减少到单个高斯均值,该均值通常落在无法到达的区域。我们提出了NFTR(具有三角松弛重加权的归一化流子目标策略)。条件归一化流取代了高斯策略,并且一个闭式模式平均结果将归一化流识别为基于AWR的子目标选择的最小生成类。基于架构三角不等式且不依赖距离准确性构建的三角松弛分数,通过乘法校正AWR权重,以降低迂回成本超过平均可达性的子目标的权重。三角松弛在确定性MDP的测地线上消失,并且在随机动力学下仍然是可组合性违反的保守上界。RWDR目标保留了AWR的总体水平单调改进,并允许进行三项次优分解。这两个要素共同产生了可证明地避免上述高斯坍塌且在随机动力学下保持稳定的子目标选择。
英文摘要
Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimistic bias treats lucky stochastic outcomes as skillful choices, and mode collapse reduces a multi-modal subgoal distribution to a single Gaussian mean that often falls in unreachable regions. We propose NFTR (Normalizing Flow subgoal policies with Triangle-slack Reweighting). A conditional Normalizing Flow replaces the Gaussian policy. A closed-form mode-averaging result identifies the Gaussian limitation, while conditional flows support multi-modal weighted maximum likelihood and direct sampling. A triangle slack score, computed from a jointly trained MRN distance head whose triangle inequality is guaranteed architecturally, corrects the AWR weight multiplicatively and measures a detour inside the learned geometry without requiring exact distance recovery. Triangle-slack vanishes on geodesics in deterministic MDPs and remains a conservative upper bound on composability violation under stochastic dynamics. The RWDR objective preserves AWR's population-level monotonic improvement and admits a three-term suboptimality decomposition. On OGBench the flow carries the larger share of the gain, while the full co-trained configuration adds task-dependent gains. The combined method avoids Gaussian mode averaging and performs well on stochastic tasks. GitHub page: https://github.com/erdemtbao/NFTR
发表机构
- Huazhong University of Science and Technology(华中科技大学)
- Xi’an Jiaotong University(西安交通大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。