arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prof-K:用于高效Top-k选择的概率单趟过滤算法

Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection

Tadeusz Dziarmaga, Witold Sikora, Łukasz Struski, Jacek Tabor, Marcin Mazur

arXiv 2608.12573首次发表:更新:

发表机构

Jagiellonian University(雅盖隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出概率单趟过滤算法Prof-K,它是与分布无关的高效Top-k选择算法,可提供概率正确性保证,在大规模场景下比PyTorch topk等方法快1.5-10倍,还能实现精度-速度权衡,可用于优化稀疏自编码器训练。

AI 中文摘要

Top-k选择是一种基础计算原语,应用范围涵盖数据库、信息检索、信号处理及现代机器学习工作负载,包括稀疏激活与注意力剪枝。随着数据规模增长,现有方法效率低下:精确方法会产生高昂的内存与计算开销,而近似方法通常依赖脆弱的启发式规则,在对抗性或重尾输入下性能会下降。本文提出Prof-K,一种快速、可扩展且与分布无关的Top-k算法,具有概率正确性保证。Prof-K执行单趟过滤流程:通过小随机样本估计自适应阈值,将N个输入元素流式传输一次至紧凑缓冲区,对该缓冲区运行精确Top-k例程,以至少1-ε(ε>0为用户指定值)的概率恢复真实Top-k元素。我们推导了正确性与缓冲区大小的高概率保证,以及作为N和k的函数最小化开销的近似最优样本量。实验表明,Prof-K相较于高度优化的PyTorch topk和最新的RadiK实现,可实现1.5倍至10倍的加速,在大规模、中小k的场景中,现有方法表现最差,而Prof-K的收益最大。与以往方法不同,这些保证独立于输入分布,确保了对抗性设置下的鲁棒性。通过放宽召回目标(例如恢复95%的真实Top-k值),Prof-K还提供了原则性的精度-速度权衡。我们进一步展示了其在训练BatchTopK稀疏自编码器(SAEs)上的影响,其中Top-k选择占训练成本的很大一部分。

英文摘要

Top-k selection is a fundamental computational primitive with applications spanning databases, information retrieval, signal processing, and modern machine learning workloads, including sparse activations and attention pruning. As data sizes grow, existing approaches become inefficient: exact methods incur high memory and compute overhead, while approximate methods often rely on brittle heuristics that degrade under adversarial or heavy-tailed inputs. In this paper, we introduce Prof-K, a fast, scalable, and distribution-agnostic algorithm for exact top-k selection. Prof-K performs a single-pass filtering procedure: a small random sample estimates an adaptive threshold, the N input elements are streamed once into a compact buffer, and an exact top-k routine on this buffer recovers the true top-k elements on the first attempt with probability at least $1-\varepsilon$, where $\varepsilon>0$ is user specified. We derive high-probability guarantees for correctness and buffer size, together with an approximately optimal sample size that minimizes overhead as a function of N and k. Empirically, Prof-K achieves 1.5x-15x speedups over the highly optimized PyTorch topk and recent RadiK implementations, with the largest gains in the large-scale, small-to-moderate-k regime where prior methods struggle most. Unlike previous approaches, these guarantees hold independently of the input distribution, ensuring robustness to adversarial settings. A run-time check detects the rare failures and triggers a retry, so the returned set is always exact and $\varepsilon$ bounds only the probability of requiring an additional pass. We further demonstrate its impact on training BatchTopK Sparse Autoencoders (SAEs), where top-k selection constitutes a significant portion of the training cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑