TokenPowerSandbox:面向能源感知大语言模型服务的证据门控CPU优先筛选
TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving
浏览论文内容
中文总结 AI 辅助
TokenPowerSandbox是面向能源感知LLM服务的证据门控CPU优先筛选工作流,结合CPU投影器、GPU探测等技术,在H100上评估Qwen2.5-7B-Instruct,验证了能源与延迟预测的关联及弃权门控的必要性。
中文摘要 AI 辅助
能源感知大语言模型(LLM)服务需要在真实请求形态下比较不同配置,但穷尽目标GPU的性能分析成本高昂,而廉价预测器在其测量范围外可能会产生危险的高置信度结果。我们提出TokenPowerSandbox,这是一种证据门控工作流,结合了可解释的驻留CPU投影器、短时间目标GPU探测、全工作负载验证以及防篡改的测量前冻结溯源。在1块NVIDIA H100 80GB GPU上,使用vLLM部署Qwen2.5-7B-Instruct模型,通过3次锚点重复和6个开发工作负载校准工作负载迁移。对同一冻结模型在盲保留集和单独预先声明的无需重新拟合的确认集上进行评估,共完成51次冻结后运行,能源平均绝对百分比误差(MAPE)分别为6.23%和7.35%,斯皮尔曼等级相关系数分别为0.976和0.933。不过,预先声明的首包时间(TTFT)门控在并发数为4时通过(MAPE为9.27%),在并发数低于4时触发弃权(不执行)(MAPE为64.80%),这表明能源精度无法保证延迟性能。
英文摘要
Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly and a cheap predictor can be dangerously confident outside its measured scope. We present TokenPowerSandbox, an evidence-gated workflow that combines an interpretable CPU-resident projector, short target-GPU probes, full-workload verification, and tamper-evident freeze-before-measurement provenance. On one NVIDIA H100 80GB serving Qwen2.5-7B-Instruct with vLLM, three anchor repeats and six development workloads calibrate workload transfer. The same frozen model is evaluated on a blind holdout and a separately predeclared no-refit confirmation totaling 51 post-freeze runs. Energy MAPE is 6.23% and 7.35%, with Spearman rank correlations of 0.976 and 0.933. However, a predeclared TTFT gate passes at concurrency four (9.27% MAPE) and triggers abstention below four (64.80%), showing why energy accuracy cannot certify latency.