AI 中文总结
该研究针对物联网终端算力不足的问题,提出延迟最优自适应拆分推理框架,通过终端执行明文前缀、边缘与云端处理FHE密文实现隐私保护,在两类数据集上取得显著加速效果且保持准确率。
AI 中文摘要
物联网(IoT)终端设备日益被要求支持隐私敏感的批量推理,但其有限的计算资源往往使得在本地完整运行卷积神经网络难以实现。本文提出了一种面向隐私保护的云-边-端协同的延迟最优自适应拆分推理框架。终端设备作为信任锚,执行明文模型前缀,使用全同态加密(FHE)对拆分后的激活值进行加密,并将密钥保存在本地,而边缘节点和云仅在FHE密文上执行分配的模型片段。我们将协同加密推理形式化为终端侧拆分点与边缘侧终止点的拆分对选择问题。所提出的规划器联合建模明文前缀执行、加密、通信、边缘侧FHE执行以及云侧FHE完成,支持卷积级和块级两种拆分粒度。在CIFAR-10和PathMNIST数据集上的实验表明,所提出的卷积级协同方案相比全云端FHE实现了约12.9倍的摊销端到端加速,相比块级方案实现了约3.9倍的加速,同时保持了对应明文模型的准确率。计入建模的通信开销后,在CIFAR-10上的摊销延迟为1033.279秒/样本,在PathMNIST上为1023.429秒/样本。
英文摘要
Internet of Things (IoT) end devices are increasingly expected to support privacy-sensitive batch inference, yet their limited computational resources often make full local execution of convolutional neural networks impractical. This paper presents a latency-optimal adaptive split inference framework for privacy-preserving cloud-edge-end collaboration. The end device acts as the trust anchor, executes the plaintext model prefix, encrypts the split activation using fully homomorphic encryption (FHE), and keeps the secret key locally, while the edge and cloud execute assigned model segments only on FHE ciphertexts. We formulate collaborative encrypted inference as a split-pair selection problem over an end-side split point and an edge-side termination point. The proposed planner jointly models plaintext prefix execution, encryption, communication, edge-side FHE execution, and cloud-side FHE completion, and supports both convolution-level and block-level split granularities. Experiments on CIFAR-10 and PathMNIST show that the proposed convolution-level collaborative scheme achieves amortized end-to-end speedups of approximately 12.9 times over full-cloud FHE and 3.9 times over the block-level alternative, while preserving the corresponding plaintext-model accuracy. Including modeled communication, the amortized latencies are 1033.279 s/sample on CIFAR-10 and 1023.429 s/sample on PathMNIST.
Comments22 pages, 3 figures, 9 tables. Accepted at ProvSec 2026