发表机构
Paderborn University(帕德博恩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CUDA-MPQS,一种全阶段在GPU上执行的自初始化二次筛法,成功分解RSA-155,创下二次筛法最大整数纪录,并在RSA-100上显著快于CPU实现。
AI 中文摘要
二次筛法(QS)是一种不规则、缓存不友好、分支繁重的整数因式分解算法;先前的GPU工作加速了它的各个阶段。我们提出了\swname{},一种自初始化二次筛法,其中\emph{每一个}阶段——多项式初始化、筛分、关系后处理、大素数匹配、\GFtwo{}矩阵构造、Block Wiedemann线性代数和平方根——都在GPU上执行,CPU仅用于编排、设置和I/O,筛分窗口99.9%的时间由GPU占用,没有设备范围或事件同步,也没有同步复制。利用它,我们在GPU上端到端地分解了512位的RSA-155挑战模数:一个64×H100筛分器为单个H100的\GFtwo{}求解提供数据,总计700.6 GPU小时和242千瓦时,其中98.4%用于筛分。据我们所知,这是二次筛法分解的最大整数,超过了之前的QS纪录RSA-150(Buhrow,YAFU,2025);自1996年以来所有可比较的通用因式分解,包括1999年RSA-155本身的分解,都使用了数域筛法。在单个设备上,我们在H100上29.2秒分解RSA-100,在消费级RTX 5070 Ti上51.0秒;在受控的相同模数测量中,双方都测量了能耗,H100比最快的CPU二次筛法YAFU(在EPYC 9655的96个核心上,106–122秒)更快,也比CADO-NFS的SIQS(292秒)和GNFS(263秒)更快。在四块独立GPU(CUDA核心数范围3.7倍)上,RSA-100的吞吐量接近与核心数成正比,内存带宽是较弱的预测因子,从RSA-100到RSA-155的七次基准因式分解的测量成本按$\LN{1/2}{1}^{1.04}$增长($R^2=0.993$)。在RSA-150上,即先前QS纪录的相同模数,我们测量到302.9 GPU小时,而那次运行用了11,664核心小时。
英文摘要
The quadratic sieve (QS) is an irregular, cache-hostile, branch-heavy integer factorization algorithm; prior GPU work accelerates individual stages of it. We present \swname{}, a self-initializing quadratic sieve in which \emph{every} stage --- polynomial initialization, sieving, relation post-processing, large-prime matching, \GFtwo{} matrix construction, Block Wiedemann linear algebra and the square root --- executes on the GPU, the CPU restricted to orchestration, setup and I/O, the sieve window 99.9\% GPU-busy with no device-wide or event synchronization and no synchronous copy. With it we factored the 512-bit RSA-155 challenge modulus end to end on GPUs: a 64$\times$H100 sieve feeding a single-H100 \GFtwo{} solve, 700.6~GPU-h and 242~kWh in total, 98.4\% of it sieving. To our knowledge this is the largest integer factored by the quadratic sieve, exceeding the prior QS record RSA-150 (Buhrow, YAFU, 2025); all comparable general factorizations since 1996, RSA-155's own in 1999 included, used the number field sieve. On a single device we factor RSA-100 in 29.2~s on an H100 and 51.0~s on a consumer RTX~5070~Ti; in a controlled same-modulus measurement, with energy measured on both sides, the H100 is faster than the fastest CPU quadratic sieve, YAFU, on 96 cores of an EPYC~9655 (106--122~s), and than CADO-NFS's SIQS (292~s) and GNFS (263~s). Across four discrete GPUs (a 3.7$\times$ range in CUDA cores) RSA-100 throughput is close to proportional to core count, memory bandwidth being a weaker predictor, and measured cost over seven benchmark factorizations from RSA-100 to RSA-155 grows as $\LN{1/2}{1}^{1.04}$ ($R^2=0.993$). On RSA-150, the identical modulus on which the prior QS record was set, we measure 302.9~GPU-h against that run's 11{,}664 core-hours.