arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同态优势算子:在全同态加密约束下稳定强化学习

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

arXiv 2610.02074首次发表:更新:

发表机构

El Oued University; King Fahd University of Petroleum and Minerals (KFUPM); Interdisciplinary Research Center For Smart Mobility and Logistics, KFUPM(瓦德大学; 法赫德国王石油矿产大学; 法赫德国王石油矿产大学智能交通与物流跨学科研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全同态加密下强化学习多项式近似发散问题,提出同态优势算子,通过零均值投影消除贝尔曼漂移,在多种实验中实现零边界违反并提升策略准确率。

AI 中文摘要

隐私保护机器学习在云端部署智能系统并处理机密数据时面临重大挑战。全同态加密(FHE)为安全计算提供了一种引人注目的解决方案,可在云计算的整个过程中保持数据的机密性。然而,将FHE应用于强化学习(RL)需要将非线性操作替换为多项式近似,而由于一种独特的递归误差现象(称为贝尔曼漂移),这些近似会灾难性地发散。本文介绍了同态优势算子(HAO),这是一种稳定框架,旨在防止基于FHE的深度强化学习中多项式近似的发散。HAO将基于优势的价值估计中的零均值中心化投影直接应用于时序差分(TD)目标。这种线性投影消除了驱动贝尔曼漂移的均匀状态值基线,在保持每个状态动作排序的同时,不需要额外的非线性乘法深度,也避免了昂贵的密文自举。所提出的HAO框架通过三层实验方法进行了评估,包括表格型马尔可夫决策过程(MDP)、使用真实CKKS加密操作的加密CartPole环境,以及具有密集连续特征的20节点物流路径规划基准。结果表明,所提出的HAO严格将网络预激活限制在安全的多项式近似域内。所提出的HAO强化学习智能体在所有使用的随机种子下实现了0%的边界违反,而仅使用正则化(L2权重衰减和梯度裁剪)在5个种子中的3个上违反了边界,未稳定的基线在83.8%的回合中违反了边界。最后,HAO智能体在表格域中将最优策略准确率提高了18.0个百分点,并且在向裁剪梯度添加DP-SGD风格的高斯噪声时保持稳定。

英文摘要

Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑