arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过哈达玛过参数化实现的稀疏高斯混合模型Q函数用于在线强化学习

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

Minh Vu, Konstantinos Slavakis

arXiv 2607.23474首次发表:更新:

发表机构

Institute of Science Tokyo; Department of Information and Communications Engineering(东京科学技术学院; 信息与通信工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究为强化学习开发在线离策略策略迭代框架,基于稀疏高斯混合模型Q函数,借助哈达玛过参数化引入,通过在线梯度下降学习几何角色,数值测试显示其在参数效率、改进速度及低参数泛化性上表现出色。

AI 中文摘要

本文基于稀疏高斯混合模型Q函数(S-GMM-QFs)开发了一种用于强化学习(RL)的在线离策略策略迭代框架。该框架通过经验回放处理分布不匹配的同时,将流数据、非平稳数据与参数空间的黎曼结构相协调。S-GMM-QFs通过哈达玛过参数化引入,通过平滑正则化实现可解释的稀疏化,便于基于黎曼的优化。过参数化使框架能从大的初始池中自适应识别有意义的组件,产生稀疏模型,其可解释性自然源于几何。通过在笛卡尔积黎曼流形上对平滑目标进行在线梯度下降来学习这些几何角色。数值测试表明,S-GMM-QFs在使用更少参数且每次观察到的转换实现更快改进的情况下,能匹配或超越深度RL方法。值得注意的是,在稀疏深度RL方法退化的低参数 regime中,参数效率和可解释性相结合以保持强泛化性。

英文摘要

This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure of the parameter space while handling distributional mismatch through experience replay. S-GMM-QFs are introduced via Hadamard overparametrization, enabling interpretable sparsification through smooth regularization that facilitates Riemannian-based optimization. Overparametrization allows the framework to adaptively identify meaningful components from a large initial pool, yielding sparse models where interpretability emerges naturally from geometry: each component's parameters (means and covariances) explicitly encode its geometric role in the ambient state-action space. These geometric roles are learned through online gradient descent on a smooth objective over a (Cartesian-product) Riemannian manifold. Numerical tests demonstrate that S-GMM-QFs match or exceed deep RL methods while using substantially fewer parameters and achieving faster improvement per observed transition. Notably, parameter efficiency and interpretability combine to maintain strong generalization in low-parameter regimes where sparsified deep RL approaches degrade.

CommentsThis work has been submitted to the Elsevier for possible publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑