arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CipherGenome:基因组混合专家模型的同态推理

CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts

Guang Yang, Fengchen Liu

arXiv 2609.35883首次发表:更新:

发表机构

University of California, Los Angeles; University of California, Berkeley(加州大学洛杉矶分校; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CipherGenome通过模块LWE加密将95.8%的专家参数外包至不可信GPU,在保持15.1B MoE基因组模型推理精度的同时,抵御反演攻击并显著降低延迟与通信开销。

AI 中文摘要

基因组基础模型正发展为稀疏混合专家(MoE)网络,其专家权重已无法容纳在持有序列的机器上,然而将私有基因组发送至租用的加速器会使其暴露:我们证明,仅托管一个专家的服务器即可恢复输入核苷酸,其top-1准确率达99.8%。我们提出CipherGenome协议,该协议将15.1B参数的MoE基因组模型的嵌入层、注意力层和路由层保留在可信的瘦客户端上,并将每个专家投影(占参数的95.8%)外包给不受信任且可能串通的GPU服务器,使用模块LWE加密。该设计利用了三个结构事实:专家层位于两个SwiGLU门之间且为线性,专家权重是公开的,GPU整数张量核心可在单次GEMM中精确计算模2^48的密文-权重乘积。客户端精确计算非线性部分并用新密钥重新加密,因此无需多项式近似或自举。在来自12个细菌基因组的72个窗口上,加密使每个token的KL散度增加2.54×10^-4 nats(95%置信区间上限为3.95×10^-4),低于预注册的非劣效性界,且与bf16推理无法区分,而相同的反演攻击降至随机水平。可复用的公共提示将端到端延迟降低3.54倍,线压缩将流量减少6.8倍,逐层填充将路由泄漏从54.9%的准确率降至8.9%,且兼容HE的int4专家与明文对应物相比仍非劣效。每个专家和token,服务器端成本比CKKS基线低六个数量级以上。

英文摘要

Genome foundation models are growing into sparse mixture-of-experts (MoE) networks whose expert weights no longer fit on the machines that hold the sequences, yet sending a private genome to rented accelerators exposes it: we show that a single server hosting one expert recovers the input nucleotides with 99.8% top-1 accuracy. We present CipherGenome, a protocol that keeps the embedding, attention and router of a 15.1B-parameter MoE genome model on a trusted thin client and outsources every expert projection, 95.8% of the parameters, to untrusted and possibly colluding GPU servers under module-LWE encryption. The design exploits three structural facts: expert layers are linear between two SwiGLU gates, expert weights are public, and GPU integer tensor cores can evaluate a ciphertext-weight product exactly modulo $2^{48}$ in a single GEMM. The client evaluates the nonlinearity exactly and re-encrypts with fresh secrets, so no polynomial approximation or bootstrapping is ever needed. On 72 windows from 12 bacterial genomes, encryption adds $2.54 \times 10^{-4}$ nats per token of KL divergence (95% CI upper bound $3.95 \times 10^{-4}$), below a pre-registered non-inferiority margin and indistinguishable from bf16 inference, while the same inversion attack falls to chance level. A reusable public hint cuts end-to-end latency by 3.54 times, wire compression reduces traffic 6.8 times, per-layer padding reduces routing leakage from 54.9% to 8.9% accuracy, and HE-compatible int4 experts remain non-inferior to their plaintext counterparts. Per expert and token, the server-side cost is more than six orders of magnitude below a CKKS baseline.

CommentsWithdrawn by the authors pending an institutional intellectual property review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑