arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向AI治理的基于物理侧信道的 workload 识别

Workload Identification with Physical Side Channels for AI Governance

Simone Gargiulo, Gabriel Kulp

arXiv 2609.00309首次发表:更新:

发表机构

Pivotal Research; Intelligence Security Laboratories(关键研究机构; 情报安全实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种基于NVIDIA H200功耗物理侧信道的工作负载识别方法,可区分AI训练、推理与非AI计算,经强化后能有效抵御对抗规避策略,为AI治理提供计算验证手段。

AI 中文摘要

AI计算验证是国际AI治理政策中首批切实可行的切入点之一。要判断前沿实验室或任何运营方是否遵守协议,监管机构需要明确其计算资源的使用情况。AI计算的基本单元是GPU,其执行的任何活动都会留下物理痕迹。本文表明,外部观察者可通过NVIDIA H200的功耗识别运行中的工作负载类别,这种物理信道与可被欺骗或重放的片上NVML遥测不同,原则上可独立于运营方合作进行观测。我们以约10 MHz的频率记录了930条5秒时长的轨迹,涵盖17个开源大语言模型(LLM)家族和25种非AI工作负载。基于该语料库,我们在训练时未见过的模型家族上进行评估,将训练与推理、非AI计算区分的准确率达97%,宏平均F1分数为0.955。AI工作负载的频谱内容主要低于约20 kHz,训练尤其可通过内存绑定的优化器更新被识别。随后将GPU运营方视为可重塑物理计算本身的对抗者,测试了4种将训练伪装成推理的规避策略,生成了额外680条对抗轨迹。针对规避策略进行强化的检测器(测试策略被留出),对4种策略中的3种的训练检测率≥99%;第4种为稀释低秩适配(LoRA),强化分类器的检测率为48%-88%,增加一条救助规则后提升至≥98%。虽然这些攻击并非对对抗行为的全面评估,但它们提供了超出真实活动的初步见解,以及用于开发和测试更强规避机制的数据集。

英文摘要

AI compute verification is one of the first tangible and tractable points for international policy aimed at AI governance. Determining whether frontier labs, or any operator, comply with agreements requires the regulating authority to discern how their compute is used. The elementary building block of AI compute is the GPU, and any activity it executes leaves a physical trace. Here, we show that an external observer can identify the class of the workload running on an NVIDIA H200 from its power draw. Unlike on-chip NVML telemetry, which can be spoofed or replayed, such a physical channel can in principle be observed independently of operator cooperation. We recorded $930$ five-second traces at $\sim 10$ MHz, covering seventeen open LLM families and twenty-five non-AI workloads. Over this corpus we separate training from inference and from non-AI computation with an accuracy of $97\%$ and a macro-averaged F1 score of $0.955$, evaluated on model families unseen during training. AI workload spectral content predominantly lies below $\sim 20$kHz and training is particularly recognizable through the memory-bound optimizer update. The GPU operator is then treated as adversarial and able to reshape the physical computation itself. Four evasion strategies are tested to disguise training as inference, producing an additional 680 adversarial traces. A detector hardened against evasion strategies, with the tested strategy held out, catches training $\geq 99\%$ of the time for three of the four strategies. The fourth, diluted low-rank adaptation (LoRA), is detected $48$--$88\%$ of the time with a hardened classifier, rising to $\geq 98\%$ with an additional rescue rule. While these attacks are not a comprehensive evaluation against adversarial behaviour, they offer initial insights beyond genuine activities and a dataset for developing and testing stronger evasion mechanisms.

Comments10 pages, 2 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑