arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35882cs.CRcs.LGq-bio.GN

GenomeOcean Anywhere:基因组MoE的私有WebGPU推理

GenomeOcean Anywhere: Private WebGPU Inference for Genome MoEs

Guang Yang, Fengchen Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出GenomeOcean Anywhere系统,利用WebGPU和拉格朗日编码计算,在浏览器上分布式运行150亿参数基因组MoE模型,实现私有推理,保持数值精度并防止单设备泄露序列。

中文摘要 AI 辅助

基因组基础模型在序列生成的地方最为有用,然而最大的模型需要数据中心加速器以及一个发送私有DNA的地方。我们探究一个150亿参数的基因组混合专家(MoE)模型是否可以在志愿者的网页浏览器上运行,专家分布在许多不可信设备上,同时不改变其预测,也不向任何单一设备泄露序列。我们构建了一个系统,其中可信协调器运行注意力和路由,而浏览器工作者通过手写的WebGPU内核运行每个专家前馈网络,并用实值拉格朗日编码计算保护专家输入:每个工作者仅接收一个高斯填充的份额,计算专家的线性映射,协调器从三个工作者中的任意两个解码。在GenomeOcean-MoE(8个专家,top-2路由,24层)上,浏览器路径在每个量化级别上都与原生PyTorch匹配,分布式路径保持在BF16数值噪声底(每token 0.0036 nats的KL散度),一个未拟合的延迟模型在模拟广域网链接下预测解码时间中位数误差在0.74%以内。我们首先展示明文专家输入并非私有:一个探针在每个深度从单个向量恢复token,一个工作者可以从300个无序token中以92%的准确率识别源基因组。使用编码专家,一个基于份额训练的自适应攻击者降至最频繁token基线,每个工作者关于每个token的信息在每个前向传递中被限制在1比特以下,保真度成本保持在BF16噪声底以下;在Chrome中,编码解码每token运行220至376毫秒,取决于路由隐藏的程度,并且当一个工作者失败时无需副本即可继续运行。

英文摘要

Genome foundation models are most useful where sequences are generated, yet the largest models need datacenter accelerators and a place to send private DNA. We ask whether a 15-billion-parameter genome mixture-of-experts (MoE) model can instead run on volunteers' web browsers, with the experts spread across many untrusted devices, without changing its predictions and without revealing the sequence to any single device. We build a system in which a trusted coordinator runs attention and routing while browser workers run every expert feed-forward network through hand-written WebGPU kernels, and we protect the expert inputs with real-valued Lagrange coded computing: each worker receives only a Gaussian-padded share, computes the expert's linear maps, and the coordinator decodes from any two of three workers. On GenomeOcean-MoE (8 experts, top-2 routing, 24 layers), the browser path matches native llama.cpp at every quantization level, the distributed path stays at the BF16 numerical noise floor (KL 0.0036 nats per token), and an unfitted latency model predicts decode time within 0.74% (median) under emulated wide-area links. We first show that plaintext expert inputs are not private: a probe recovers the token from a single vector at every depth, and one worker can identify the source genome from 300 unordered tokens with 92% accuracy. With coded experts, an adaptive attacker trained on shares falls to the most-frequent-token baseline, one worker's information about each token is bounded below one bit per forward pass, and the fidelity cost stays below the BF16 noise floor; in Chrome, coded decoding runs at 220 to 376 ms per token, depending on how much of the routing is hidden, and continues without replicas when a worker fails.

发表机构

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of California, Berkeley(加利福尼亚大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑