YAVIN:面向安全边缘内存处理的统一架构
YAVIN: A Unified Architecture for Secure Edge Processing in Memory
浏览论文内容
中文总结 AI 辅助
针对边缘计算中多租户安全执行及内存处理的挑战,提出将TEE扩展至内存的YAVIN架构,实现后量子密码与认证加密的PIM协同设计,在量化LLM执行时获超20倍加速且开销极低。
中文摘要 AI 辅助
安全、私密的多租户执行(涵盖处理器、内存和加速器)仍是现代边缘计算系统面临的最重大挑战之一。与此同时,内存处理(Processing-in-Memory,PIM)作为一种将计算移近数据以缓解冯·诺依曼瓶颈的有效方法已应运而生。现有的可信执行环境(Trusted Execution Environments,TEE)仅在处理器内部建立信任,在数据通过内存总线等不可信资源时提供保护,因此无法在内存中直接执行可信计算。本文提出YAVIN,一种统一的可信计算基(Trusted Computing Base,TCB),它将TEE的信任范围扩展至处理器之外,涵盖处理器执行和支持可信内存处理执行的专用内存区域,同时将内存总线视为不可信资源。YAVIN利用传统TEE架构已建立的专用受保护内存区域,使数据在TEE内可由处理器或PIM执行完成解密、处理和重新加密。为实现这一统一TCB,YAVIN实现了LightSaber密钥 Encapsulation Mechanism(Key Encapsulation Mechanism,KEM)后量子密码系统和ASCON-128认证加密的首个PIM实现,通过协同设计两种算法以实现高效DRAM执行,从而建立并维护共享密码状态。最后,本文展示了面向张量工作负载的密码学-PIM协同设计如何重组计算,以满足认证加密施加的排序约束,同时将临时明文暴露限制为位片排序,且仅产生最小性能开销。与最新的PIM AES实现相比,YAVIN实现了超过20倍的加速,在执行INT8和INT32量化边缘级大语言模型(LLM)时,相对于明文执行仅分别产生34%和9.3%的开销。
英文摘要
Secure, private multi-tenant execution spanning processors, memory, and accelerators remains one of the most significant challenges in modern edge computing systems. Simultaneously, processing-in-memory (PIM) has emerged as an effective approach for reducing the Von Neumann bottleneck by moving computation closer to data. Existing trusted execution environments (TEEs) establish trust only within the processor, protecting data while it traverses untrusted resources such as the memory bus. Consequently, trusted computation cannot be performed directly within memory. We present YAVIN, a unified trusted computing base (TCB) that extends the TEE beyond the processor to encompass both processor execution and a dedicated memory region supporting trusted processing-in-memory execution while treating the memory bus as untrusted. Leveraging the dedicated protected memory regions already established by conventional TEE architectures, YAVIN enables data to be decrypted, processed, and re-encrypted by either processor or PIM execution while remaining within the TEE. To realize this unified TCB, YAVIN presents the first PIM implementations of the LightSaber KEM post-quantum cryptosystem and ASCON-128 authenticated encryption, co-designing both algorithms for efficient DRAM execution to establish and maintain shared cryptographic state. Finally, we demonstrate how cryptography-PIM co-design for tensor-based workloads reorganizes computation to satisfy the ordering constraints imposed by authenticated encryption with minimal performance overhead while simultaneously enabling bit-sliced ordering that limits temporary plaintext exposure. Compared to the latest PIM AES implementation, YAVIN achieves more than a 20x speedup while incurring only 34% and 9.3% overhead when executing INT8 and INT32 quantized edge-class LLMs, respectively, relative to plaintext execution.