arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06912cs.AI

Fast LapSum:百万级规模下的精确可微Top-k操作

Fast LapSum: Exact Differentiable Top-$k$ at Million Scale

  • Faculty of Mathematics and Computer Science, Jagiellonian University(雅盖隆大学数学与计算机科学学院)
  • Wrocław University of Science and Technology(弗罗茨瓦夫理工大学)
  • Centre for Credible Artificial Intelligence, Warsaw University of Technology(华沙理工大学可信人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

Jakub Antczak, Joanna Wojciechowicz, Kamil Książek, Marcin Mazur, Łukasz Struski, Jacek Tabor

AI总结:

研究针对大规模模型中现有软Top-k方法成本过高的问题,提出Fast LapSum,该方法可在百万级规模下实现精确可微的软Top-k,在两个应用中展现出高效性与先进性。

AI中文摘要:

Top-k操作是现代稀疏计算的基础模块,可用于token路由、专家激活、内存选择和注意力剪枝。然而,标准硬Top-k操作会阻断梯度,而现有的连续(软)松弛方法对于大规模模型来说成本过高。我们提出Fast LapSum,这是一种精确预算的软Top-k原语,其GPU求解器在排序后以线性时间运行。与之前的线性时间方法(如DFTopK,其放松了归一化约束)不同,据我们所知,Fast LapSum是第一种在保持k的精确选择质量的同时,仍能完全端到端可微的方法。我们的求解器结合了线性时间阈值计算与解析向量-雅可比乘积,对于极端规模,采用概率 bracketing 仅对核噪声分数的不确定中间带进行排序。由此产生的开销几乎可以忽略不计:求解器处理10^6、10^7和10^8个分数分别耗时0.41、1.15和5.23毫秒。这使得精确软Top-k在稀疏路由、检索和大规模优化中变得实用。我们在训练循环中处理数百万坐标的两个高要求应用上展示了Fast LapSum:生成百万像素级稀疏对抗样本,其精确软预算约为图像像素的0.02%,实现了比现有最先进方法快一个数量级的速度;以及从头开始训练完全可微的稀疏图像编码器。

英文摘要:

Selecting the top-$k$ elements is a fundamental operation for inducing sparsity in large-scale models and optimization problems, enabling robust expert activation, token routing or attention pruning. However, hard top-$k$ is non-differentiable, while existing differentiable alternatives become increasingly expensive as the number of coordinates grows. We introduce Fast LapSum, a scalable solver for the LapSum soft top-$k$ formulation that preserves an exact selection mass of $k$, while supporting end-to-end differentiation. In Fast LapSum, we reduce sorting cost using probabilistic bracketing, which restricts sorting to a narrow band of scores around the threshold using a binomial order-statistic from kernel-noised samples. A certification pass upgrades the probabilistic localization to a verified one at the cost of one additional linear pass, with a full-sort fallback that covers the worst-case scenario. Our certified GPU implementation processes $10^6$, $10^7$, and $10^8$ scores in median times of $0.92$, $1.39$, and $7.24$ ms, respectively, making exact-budget soft top-$k$ practical within million-scale optimization loops. We demonstrate this capability in two applications: megapixel sparse adversarial examples with a small fraction of initial image pixels, where Fast LapSum achieves an order-of-magnitude speedup over state-of-the-art methods, and 3D Gaussian splatting. In the latter, we use Fast LapSum to reduce the number of Gaussians produced by Adaptive Density Control to a substantially smaller number while retaining nearly the same rendering quality and massively reducing computation. These results demonstrate that exact-budget differentiable top-$k$ can be incorporated into practical million-scale optimization pipelines.

↑