arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17442cs.CR

FESC:基于加密状态空间模型重构长上下文私有推理

FESC: Remodeling Long-Context Private Inference with Encrypted State-Space Models

Yufan Zhu, Chao Jin, Khin Mi Mi Aung, Xiaokui Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

FESC是首个可在单GPU上完成L≥1024原生端到端执行的私有长文档推理系统,通过混合FHE-MPC架构优化选择性SSM推理,在L=2048时12层Mamba-base模型推理耗时77.3分钟且准确率接近明文。

中文摘要 AI 辅助

使用机器学习模型处理长且敏感的文档需要高效且保护隐私的长上下文推理。现有的私有推理系统会优化或分布式加密Transformer注意力机制,但随着序列长度增长,其二次型的 token 对计算仍是瓶颈。选择性状态空间模型(SSMs)提供线性时间递归,但直接加密实现会产生线性乘法深度、全序列状态驻留或密集FHE-MPC转换。我们提出Factorized Encrypted Scan-Contract(FESC),一种用于私有长上下文选择性SSM推理的混合FHE-MPC系统。其因子化解码-收缩在转换边界间保持依赖输入的转换紧凑性,无需密集扩展即可组合,按需流式传输状态块,并在转换前收缩输出。我们验证了该扫描-收缩实现在不变型和选择性SSM架构间的接口兼容性。针对Mamba-2实例,我们设计了用于线性计算的GPU优化CKKS核、用于SiLU、softplus、指数和RMSNorm的MPC协议,并采用感知近似的微调。据我们所知,FESC是首个在单GPU上完成L≥1024时原生端到端执行的私有长文档推理系统。当L=2048时,12层Mamba-base模型在单个A100 GPU上完成推理耗时77.3分钟,峰值内存占用为32.7 GB,同时在评估的长文档任务上保持接近明文的准确率。

英文摘要

Processing long, sensitive documents with machine-learning models requires efficient, privacy-preserving long-context inference. Prior private inference systems optimize or distribute encrypted Transformer attention, but its quadratic token-pair work remains the bottleneck as sequence length grows. Selective state-space models (SSMs) offer linear-time recurrence, yet direct encrypted implementation incurs linear multiplicative depth, sequence-wide state residency, or dense FHE-MPC conversion. We present Factorized Encrypted Scan-Contract (FESC), a hybrid FHE-MPC system for private long-context selective SSM inference. Its factorized scan-contract keeps input-dependent transitions compact across conversion boundaries, composes them without dense expansion, streams state chunks on demand, and contracts outputs before conversion. We demonstrate interface compatibility of the scan-contract implementation across invariant and selective SSM architectures. For our Mamba-2 instantiation, we design GPU-optimized CKKS kernels for linear computations, MPC protocols for SiLU, softplus, exponential, and RMSNorm, with approximation-aware fine-tuning. To our knowledge, FESC is the first private long-document inference system to complete native end-to-end execution at $L \geq 1{,}024$ on a single GPU. At $L = 2{,}048$, a 12-layer Mamba-base model completes inference in 77.3 minutes on one A100 GPU with a peak memory footprint of 32.7 GB, while maintaining near-plaintext accuracy on the evaluated long-document tasks.

补充信息

↑