arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18069cs.ARcs.LG

动态限制人工智能性能的硬件机制

Hardware Mechanisms to Dynamically Throttle AI Performance

Haiyue Ma, Lauren Malek, Joseph Forzani, David Wentzlaff

首次发表
浏览论文内容

中文总结 AI 辅助

研究在人工智能模型集成到关键系统时缺乏意图控制的问题,提出用微架构旋钮动态控制硬件资源限制AI性能,评估候选旋钮并构建机制,展示其高性能敏感度等优势及多旋钮组合效果。

中文摘要 AI 辅助

随着功能更强大的人工智能模型越来越多地集成到关键计算机系统中,缺乏对人工智能意图的控制促使人们建立安全机制。现有软件保护措施仅施加行为约束,可能被足够智能的模型绕过。虽然硬件级安全强制被视为至关重要的最后一道防线,但除了对未经授权访问的政策规定或粗略的全芯片关闭外,很少有机制被提出。本文引入了一组微架构旋钮,可在运行时动态控制可用硬件资源以限制人工智能性能。评估了跨越GPU内存子系统在容量、带宽、延迟和频率维度的候选旋钮,确定了四个有力候选者:L2大小、L2延迟、L2带宽和共享内存端口访问速率。为最小化新逻辑和额外设计成本,用成熟的微架构原语构建了所有四种机制。结果表明这些旋钮具有高性能敏感度、可忽略的实现成本、动态节流后快速稳定以及对芯片其余部分的最小附带影响。此外,多旋钮分析揭示了能放大性能下降的旋钮组合,可实现更广泛的性能目标。

英文摘要

As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. Existing software safeguards impose only behavioral constraints that can potentially be bypassed by sufficiently intelligent models. While hardware-level safety enforcement has been recognized as an essential last line of defense, few mechanisms have been proposed beyond policy regulations on unauthorized accesses or coarse full-chip shutdown. What is missing is a fine-grained, dynamic intervention mechanism at the architecture level. In this paper, we introduce a set of microarchitecture knobs which dynamically control the available hardware resources to limit AI performance at runtime. We evaluate candidate knobs spanning the GPU memory subsystem, across capacity, bandwidth, latency and frequency dimensions, and narrow down to four strong candidates: L2 size, L2 latency, L2 bandwidth, and shared memory port access rate. To minimize new logic and extra design cost, we build all four mechanisms from well-established microarchitectural primitives: cache way masking, credit-based rate limiting, latency insertion, and bank arbitration. We show that these knobs achieve high performance sensitivity (up to 80% performance cut at 1/8 resource availability), negligible implementation cost (<~10K flip flops), fast stabilization after dynamic throttling (5-80K cycles), and minimal collateral impact on the rest of the chip. Further, multi-knob analysis reveals combinations of knobs that amplify the performance degradation beyond the effect of each knob individually, which enables a broader range of performance targets.

发表机构

  • Princeton University(普林斯顿大学)
  • Princeton University Electrical(普林斯顿大学电子与计算机工程)

机构由 AI 辅助整理,请以论文原文为准。

↑