发表机构
Columbia University; UC Berkeley(哥伦比亚大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分布式推理管道中网络驱动的精度崩溃问题,通过模拟自动驾驶中的两层边缘云多目标跟踪管道,发现shaped workload攻击可致精度崩溃,如使良性p99延迟增加,降低跟踪质量,促使相关攻击与防御研究。
AI 中文摘要
推理系统越来越多地将在应用程序延迟期限内返回预测的快速路径与在更强的远程硬件上运行更高计算方法的高精度慢速路径相结合,以便及时返回结果并与快速路径预测相结合。在这项工作中,我们表明这种新的协调层暴露了一个新的攻击面:例如,Yo-Yo突发等 shaped workload攻击可以利用慢速路径上共享资源的争用,将良性用户的慢速路径预测推过其延迟期限。合并器随后会丢弃这些预测,而快速路径继续返回及时的输出。我们将由此导致的慢速路径精度优势的丧失称为精度崩溃。我们在自动驾驶的两层边缘云多目标跟踪管道中演示了精度崩溃。在模拟中,大约4000个突发形状的请求将良性p99延迟从92毫秒增加到2秒,几乎消除了慢速路径云推理的优势,平均将目标跟踪质量降低了7.0个HOTA点。我们进一步发现,精度下降会因攻击所针对的视频间隔而有很大差异(2.0-18.7个HOTA点),并且某些罕见类别(例如停车标志)的攻击前预测精度损失近一半。这些结果表明,workload攻击可以在无需访问模型权重或受害者数据的情况下降低预测质量,并促使人们对这些新兴推理管道架构中的路由、合并、调度和资源隔离的攻击和防御进行研究。
英文摘要
Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions. Across several application domains, we abstract this inference architecture as a fast path, a slow path, and a coordination layer with two functions: a router that invokes the slow path and a merger that decides whether to incorporate its returned predictions. In this work, we show that this new coordination layer exposes a new attack surface: shaped workload attacks, e.g., Yo-Yo bursts, can exploit contention at shared resources along the slow path to push benign users' slow-path predictions past their latency deadlines. The merger then discards those predictions, while the fast path continues to return timely outputs. We refer to the resulting loss of slow-path accuracy benefits as accuracy collapse. We demonstrate accuracy collapse in a two-tier edge-cloud multi-object tracking pipeline in autonomous driving. In simulation, approximately 4,000 burst-shaped requests increase benign p99 latency from 92ms to 2s, nearly eliminating the benefit of the slow path's cloud inference, reducing object tracking quality by 7.0 HOTA points on average. We further find that accuracy degradation can significantly vary (2.0-18.7 HOTA points), depending on the video intervals that are targeted in the attack, and that certain rare classes (e.g., stop signs) lose nearly half of their pre-attack prediction accuracy. These results show that workload attacks can degrade prediction quality without needing either access to model weights or victim data, and motivate research on attacks and defenses for routing, merging, scheduling, and resource isolation in these emerging inference pipeline architectures.