发表机构
Moscow Independent Research Institute of Artificial Intelligence; Mohamed bin Zayed University of Artificial Intelligence; Lomonosov Moscow State University; Innopolis University(莫斯科独立人工智能研究所; 穆罕默德·本·扎耶德人工智能大学; 莫斯科罗蒙诺索夫国立大学; 因诺波利斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究熵正则化线性规划,提出一种加速随机块方法,利用熵特定Hessian间隙界保持稀疏数据访问成本,并比较两种对偶表示下的算法性能。
AI 中文摘要
我们研究了具有稀疏仿射约束的熵正则化线性规划,并比较了由是否保留或消除冗余归一化约束所引发的两种精确对偶表示。消除该约束会产生一个部分可分离的 sum-exp 对偶,其具有稀疏仿射因子但曲率无界。我们的主要结果是一种加速的随机块方法,其加速保持了稀疏数据访问成本。关键要素是一个全局的熵特定 Hessian 间隙界,该界对带符号的约束矩阵仍然有效,允许可计算的块重叠细化,并在采样前提供确定性曲率证书。这种几何结构支持平方根重要性采样和一种精确的惰性实现,其中稀疏算术成本由局部曲率而非最差块因子加权。我们进一步引入了一个依赖于间隙的曲率维度,将广义光滑机制与坐标方法的经典有效秩直觉联系起来。保留归一化约束会产生一个具有全局有界高阶导数的 log-sum-exp 对偶,从而支持梯度正则化和三次正则化牛顿方法。实验从坐标轮次、算术运算和墙钟时间方面比较了所得方法,分离了加速、触及数据稀疏性和问题特定结构的影响。
英文摘要
We study entropy-regularized linear programs with sparse affine constraints and compare two exact dual representations induced by whether a redundant normalization constraint is retained or eliminated. Eliminating it yields a partially separable sum-exp dual with sparse affine factors but unbounded curvature. Our main result is an accelerated randomized block method whose acceleration preserves sparse data-access cost. The key ingredient is a global entropy-specific Hessian--gap bound that remains valid for signed constraint matrices, admits computable block-overlap refinements, and provides deterministic curvature certificates before sampling. This geometry supports square-root importance sampling and an exact lazy implementation in which sparse arithmetic cost is weighted by local curvature rather than by a worst-block factor. We further introduce a gap-dependent curvature dimension that connects the generalized-smooth regime to the classical effective-rank intuition for coordinate methods. Retaining the normalization constraint yields a log-sum-exp dual with globally bounded higher derivatives, enabling gradient-regularized and cubic-regularized Newton methods. Experiments compare the resulting methods in terms of coordinate epochs, arithmetic operations, and wall-clock time, separating the effects of acceleration, touched-data sparsity, and problem-specific structure.