arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21364cs.CR

基于批处理的Paillier同态加密二进制嵌入汉明距离计算

Batched Paillier-Based Hamming-Distance Computation over Binary Embeddings

Yavor Litchev, Liwen Ouyang

首次发表
浏览论文内容

中文总结 AI 辅助

针对加密二进制嵌入的汉明距离计算,提出结合编码优化、预计算和GPU批处理的Paillier方案,实现较CPU基线数百倍的吞吐量提升。

中文摘要 AI 辅助

加法同态加密支持对加密的二进制嵌入进行外包计算,但大整数运算和数据移动可能限制吞吐量。我们描述了一个基于Paillier的客户端,它结合了进位分离的二进制编码、基于表的加密、降指数解密、CUDA/CGBN算术、持久设备状态和批处理检索集成。我们建立了该编码的正确性,并描述了四种CPU和GPU客户端配置。查找表配置使用280位指数大小参数。在3次热状态试验中,对10,000个随机512位嵌入的批次,查找表GPU配置实现了中位批次吞吐量43,091次加密/秒和28,983次汉明距离解码/秒。其分摊成本为每向量0.0232毫秒和0.0345毫秒,相对于测量的CPU基线分别对应453.8倍和200.9倍的提升。这些实现特定的结果展示了结合密码学预计算、批处理加速器执行和持久运行时状态的吞吐量优势。该研究区分了热批次性能与孤立请求延迟,并识别了初始化、传输和检索集成的剩余成本。

英文摘要

Additively homomorphic encryption supports outsourced computation on encrypted binary embeddings, but large-integer arithmetic and data movement can limit throughput. We describe a Paillier-based client that combines a carry-separated binary encoding, table-based encryption, reduced-exponent decryption, CUDA/CGBN arithmetic, persistent device state, and batched retrieval integration. We establish the encoding's correctness and characterize four CPU and GPU client configurations. The lookup configuration uses a 280-bit exponent-size parameter. Across 3 warm-state trials on batches of 10,000 random 512-bit embeddings, the lookup GPU configuration achieved median-batch throughputs of 43,091 encryptions/s and 28,983 Hamming- distance decodes/s. Its amortized costs were 0.0232 ms and 0.0345 ms per vector, corresponding to factors of 453.8 and 200.9 relative to the measured CPU baseline. These implementation- specific results demonstrate the throughput benefits of combining cryptographic precomputation, batched accelerator execution, and persistent runtime state. The study distinguishes warm-batch performance from isolated-request latency and identifies the remaining costs of initialization, transport, and retrieval integration.

↑