arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16771cs.CRcs.ARcs.ET

SABLE:用于受限机密计算的极简指令级认证加密

SABLE: Minimalist Instruction-Level Authenticated Encryption for Constrained Confidential Computing

Hamid Noori, Carlton Shepherd

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对受限机密计算,提出RISC-V处理器架构SABLE,通过在特定点集成解密验证阶段及探索多种微架构实现指令级认证加密,在FPGA上评估其性能并讨论设计权衡。

中文摘要 AI 辅助

传统处理器设计在整个执行过程中将代码和数据以明文形式暴露,易受窃取知识产权或修改安全/安全检查的攻击。指令级加密(ILE)可在运行时对单个加密程序指令进行CPU级解密、执行和可选认证。但现有方案依赖特定微架构,在指令执行后检测损坏指令,依赖非标准密码,或需要复杂的程序状态分析。本文介绍并展示了一种RISC-V处理器架构(SABLE)的设计探索,它能对程序进行微创指令级认证加密。SABLE与底层微架构无关,只需对编译后的ELF二进制文件进行少量后处理更改,就能与标准RISC-V工具链兼容。我们在两个点(指令内存包装器和CPU前端)集成了解密和验证阶段,并探索了从单周期(组合)设计到六个多周期(顺序)变体的七种ILE微架构。我们使用开源NEORV32片上系统在Xilinx Artix-7 FPGA上实现并评估了这些设计。使用Dhrystone基准测试套件,相对于基线性能,配置的LUT、性能、功耗和每条指令的能量开销分别为1.6 - 9.3倍、4.1 - 10.0倍、1.5 - 8.0倍和10.4 - 80.0倍。最后,我们讨论了设计权衡,突出了面积、性能和能量感知的设计要点。

英文摘要

Conventional processor designs expose code and data as plaintext throughout execution, rendering them inherently vulnerable to attacks that recover intellectual property or modify security/safety checks. Instruction-level encryption (ILE) enables CPU-level decryption, execution, and optionally authentication of individual encrypted program instructions at runtime. However, existing proposals depend on specific micro-architectures, detect corrupted instructions after they have executed, rely on non-standard ciphers, or require complex analyses of program state. In this work, we introduce and present a design exploration of a RISC-V processor architecture (SABLE) that enables minimally invasive instruction-level authenticated encryption of programs. SABLE is agnostic to the underlying micro-architecture, remaining compatible with the standard RISC-V toolchain with minor changes to post-process compiled ELF binaries. We integrate a decrypt-and-verify stage at two points (the instruction-memory wrapper and the CPU frontend) and explore seven ILE micro-architectures from a single-cycle (combinational) design to six multi-cycle (sequential) variants. We implement and evaluate the designs using ASCON-128a on a Xilinx Artix-7 FPGA with the open-source NEORV32 system on chip. Relative to baseline performance, the configurations span LUT, performance, power, and energy-per-instruction overheads of 1.6-9.3$\times$, 4.1-10.0$\times$, 1.5-8.0$\times$, and 10.4-80.0$\times$, respectively, using the Dhrystone benchmarking suite. Finally, we discuss design trade-offs, highlighting area-, performance-, and energy-aware design points.

↑