arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31087cs.CR

BenX:用于计算完整性的资源共享置换

BenX: Resource-Sharing Permutations for Computational Integrity

Luca Campa, Thomas De Cnudde, Al Kindi, Arnab Roy, Fabian Schmid, Markus Schofnegger, Stefano Trevisani

首次发表
浏览论文内容

中文总结 AI 辅助

BenX是一种基于Benes网络和Dickson多项式的置换哈希函数,支持硬件加速与资源共享,在FPGA上吞吐量媲美Poseidon2,在零知识证明中速度显著优于Monolith。

中文摘要 AI 辅助

基于素数模整数上的密码哈希函数在计算完整性的证明系统的效率和安全性中起着决定性作用。早期设计侧重于紧凑的算术电路和高效的软件执行,主要针对通用CPU而非硬件加速器。这项工作致力于在高效软件执行的同时,实现高效的资源共享和硬件加速。我们提出BenX,一种基于置换的哈希函数,专为硬件加速、快速CPU执行和证明系统中的低电路复杂度而设计。为了构造底层置换,我们利用Dickson多项式将Benes网络转化为$\mathbb{F}_p^2$上的可逆函数。其结构产生的置换特别适合硬件资源共享和加速。在FPGA上完全流水线化后,BenX的吞吐量与Poseidon2相当,对于Goldilocks和BabyBear,延迟分别低1.36倍和2.15倍。虽然作为独立核心较大,但它能更有效地利用共享硬件:在我们的双模架构中,NTT和哈希共享域乘法器,BenX使其中100%保持忙碌,而Poseidon和Poseidon2的部分轮次仅为6.25-20.3%。在软件方面,BenX在所有测试的31位和64位配置中均优于Poseidon,但比Poseidon2慢,除了16元素BabyBear实例。在零知识证明中,BenX证明Goldilocks置换比Monolith快6-10倍。其紧凑的算术化在BabyBear上使用的迹单元比Poseidon和Poseidon2少约36%,而其快速变体所需的证明者时间约为Poseidon2的1.3-2.6倍。

英文摘要

Cryptographic hash functions over integers modulo a prime play a decisive role in the efficiency and security of proof systems for computational integrity. Early designs focused on compact arithmetic circuits and efficient software execution, primarily targeting general-purpose CPUs rather than hardware accelerators. This work focuses on enabling efficient resource sharing and hardware acceleration alongside efficient software execution. We propose BenX, a permutation-based hash function designed for hardware acceleration, fast CPU execution, and low circuit complexity in proof systems. To construct the underlying permutation, we turn the Benes network into an invertible function over $\mathbb{F}_p^2$ using a Dickson polynomial. Its structure yields a permutation particularly suited to hardware resource sharing and acceleration. Fully pipelined on FPGA, BenX matches the throughput of Poseidon2, with 1.36x and 2.15x lower latency for Goldilocks and BabyBear, respectively. Although larger as a standalone core, it makes more effective use of shared hardware: in our dual-mode architecture, where the NTT and hash share field multipliers, BenX keeps 100% of them busy, compared with 6.25-20.3% for the partial rounds of Poseidon and Poseidon2. In software, BenX outperforms Poseidon in all tested 31- and 64-bit configurations, but is slower than Poseidon2, except for the 16-element BabyBear instance. In zero-knowledge proofs, BenX proves Goldilocks permutations 6-10x faster than Monolith. Its compact arithmetization uses about 36% fewer trace cells than Poseidon and Poseidon2 over BabyBear, while its fast variant requires 1.3-2.6x the prover time of Poseidon2.

发表机构

  • Fhenix
  • Miden
  • Technical University of Graz(格拉茨工业大学)
  • [[alloc] init]

机构由 AI 辅助整理,请以论文原文为准。

↑