arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15505math.OCcs.LGcs.NAcs.NEmath.NA

通过保扰动内存压缩实现快速且可扩展的卡普托分数阶梯度下降

Fast and Scalable Caputo Fractional Gradient Descent via Perturbation-Preserving Memory Compression

Hwanseo Lee, Junseo Lee, Hyunju Kim

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何让基于卡普托的优化计算可行,核心方法是将分数阶下降方向表示为离散卷积,引入指数和近似与二元分层离散卷积降低内存成本,贡献是在特定假设下保证方法有单调下降和线性收敛性。

中文摘要 AI 辅助

分数阶梯度下降(FGD)通过卡普托型算子纳入长程记忆,并已证明可提高病态和非凸优化问题的稳定性。但其实际应用受限,主要因评估历史依赖卷积的计算成本高,与迭代次数成二次方关系。本文致力于在不牺牲其固有内存结构的情况下使基于卡普托的优化在计算上可行。首先将分数阶下降方向表示为过去梯度的离散卷积,在此基础上引入两种互补机制降低内存项成本。第一种使用幂律核的指数和(SOE)近似实现高效递归更新,第二种新提出的二元分层离散卷积(DHDC)通过多尺度聚合策略压缩梯度历史。将这些近似视为理想卡普托算子的扰动,在标准μ强凸性和L光滑性假设下,表明只要近似误差可控,所得方法仍具有单调下降和线性收敛性。

英文摘要

Fractional gradient descent (FGD) incorporates long-range memory through Caputo-type operators and has been shown to improve stability in ill-conditioned and nonconvex optimization problems. Despite these advantages, its practical use remains limited, mainly due to the high computational cost of evaluating history-dependent convolutions, which scales quadratically with the number of iterations. In this paper, we focus on making Caputo-based optimization computationally viable without sacrificing its intrinsic memory structure. We begin by expressing the fractional descent direction as a discrete convolution over past gradients, which provides a unified view of the method. Based on this formulation, we introduce two complementary mechanisms to reduce the cost of the memory term. The first uses a sum-of-exponentials (SOE) approximation of the power-law kernel, leading to efficient recursive updates. The second approach, newly proposed in this paper as dyadic hierarchical discrete convolution (DHDC), compresses the gradient history through a multiscale aggregation strategy. Rather than treating these approximations as purely numerical accelerations, we interpret them as perturbations of the ideal Caputo operator. This viewpoint allows us to analyze how the compressed memory affects the optimization dynamics. Under standard $μ$-strong convexity and $L$-smoothness assumptions, we show that the resulting method still exhibits monotone descent and linear convergence, provided that the approximation error remains controlled.

发表机构

  • Department of Energy Engineering, Korea Institute of Energy Technology, Naju 58330, Republic of Korea(能源工程系,韩国能源技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑