arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

恶意节点下对等分布式大语言模型推理的完整性

Integrity of peer-to-peer distributed LLM inference under malicious nodes

Mert Cihangiroglu, Antonino Nocera

arXiv 2607.19490首次发表:更新:

发表机构

DCALab, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy(DCALab,电气、计算机与生物医学工程系,帕维亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究对等分布式LLM推理在恶意节点下的完整性问题,提出通过测量节点激活值变化检查输出完整性的方法,即混入秘密金丝雀输入,利用与已知参考的偏差识别恶意节点,实验中检测器AUROC达1.0。

AI 中文摘要

对等分布式推理通过在多个节点上分布层来在聚合的消费硬件上执行大语言模型(LLM)。每个请求都会经过由多个独立方拥有和控制的节点。然而,在此设置中,任何一方都可能篡改其层的输出以破坏最终结果。在可信硬件上重新计算前向传播可以发现这一点,但会引入额外的计算成本。科学文献中有几种先前的完整性检查方法,但这些解决方案仅测试精确的正确性,未考虑良性节点之间可能出现的正常变化。本文提出一种通过测量每个节点传递给下一个节点的激活值变化来检查输出完整性的方法。想要使用网络的对等方选择一小部分预先知道正确激活值的秘密金丝雀输入并将它们混入常规流量中。由于对等方无法区分金丝雀输入和真实查询,任何篡改节点也会破坏它们。与已知参考的偏差然后揭示恶意活动:良性节点仅表现出与硬件引起的噪声有微小变化,而被篡改的节点偏差要大得多。我们将恶意节点的识别视为一种概率测试,它分离两个漂移分布,而不依赖于固定阈值。我们在任何实验运行之前固定度量和成功标准来研究408种配置;检测器达到AUROC 1.0,在每种配置中的每个金丝雀上都正确地将恶意分片排在每个良性分片之上。

英文摘要

Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many nodes. Every request passes through nodes that are owned and controlled by multiple independent parties. However, in this setting, any party can tamper with the output of its layers to corrupt the end result. Recomputing the forward pass on trusted hardware can catch this, but it introduces additional computational cost. The scientific literature includes several prior integrity-checking approaches, such as known-answer traps for image classifiers and cryptographic commitments. However, these solutions test only the exact correctness and do not account for the ordinary variation that may arise between benign nodes. In this paper, we propose a method that checks the output integrity by measuring the variation in the activations that each node passes to the next. A peer who wants to use the network selects a small set of secret canary inputs whose correct activations are known in advance and mixes them into regular traffic. Because the peers cannot tell a canary from a real query, any tampering node corrupts them as well. The deviation from the known reference then reveals malicious activity: benign nodes exhibit only minor variation from hardware-induced noise, whereas tampered nodes deviate far more. We treat the identification of malicious nodes as a probabilistic test that separates two drift distributions, without relying on a fixed threshold. We study 408 configurations with metrics and success criteria fixed before any experiment ran; the detector reaches AUROC 1.0, correctly ranking the malicious shard above every benign shard on every canary in every configuration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑