arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03430cs.AI

插队:利用LLM调度中的长度预测漏洞

Jumping the Line: Exploiting Length Predictions in LLM Scheduling

  • Florida State University(佛罗里达州立大学)
  • MIT(麻省理工学院)
  • Tel Aviv University(特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuyang Dai, Rana Shahout, Mahmood Sharif

AI总结:

提出JIL攻击,通过对抗性后缀操纵长度预测,使LLM调度器低估请求长度以插队提速,实验显示预测长度最多降83.4%,完成速度提升1.53倍,并评估了防御措施。

AI中文摘要:

在大语言模型(LLM)服务中,高效的请求调度对于减少完成时间日益重要。基于大小的策略(如最短作业优先)优先处理较短的请求,但输出长度在生成前是未知的,因此实际调度器依赖预测长度。我们提出了JIL,一种针对基于预测的LLM调度器的攻击方法,通过操纵调度信号来获得更高优先级并减少完成时间。以TRAIL作为案例研究,JIL优化一个对抗性后缀,使轻量级输出长度探针低估请求的长度。我们在两个数据集和四个LLM上,针对不同的请求配置和部署配置评估了JIL。JIL将预测输出长度最多减少83.4%,在端到端服务实验中,对抗性请求平均完成速度提高最多1.53倍。预测长度的减少幅度远大于实际输出长度的变化,揭示了调度器估计与请求实际大小之间的不匹配。响应效用因模型和任务而异,暴露出调度优势与响应质量之间的权衡。我们还评估了调度器侧防御措施,发现将长度预测分组为粗略区间可降低JIL的调度优势,并减轻对良性请求的延迟。

英文摘要:

Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an attack on prediction-based LLM schedulers that manipulates the scheduling signal to obtain higher priority and reduce completion time. Using TRAIL as a case study, JIL optimizes an adversarial suffix that causes a lightweight output-length probe to underestimate a request's length. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. JIL reduces predicted output lengths by up to 83.4 percent, and adversarial requests complete up to 1.53 times faster on average in end-to-end serving experiments. The reduction in predicted length is substantially larger than the change in actual output length, revealing a mismatch between the scheduler's estimate and the request's realized size. Response utility varies across models and tasks, exposing a trade-off between scheduling advantage and response quality. We also evaluate scheduler-side defenses and find that grouping length predictions into coarse intervals reduces JIL's scheduling advantage and mitigates delays to benign requests.

补充信息

↑