arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SlackDrive:回收运行时松弛以实现自适应驾驶推理

SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference

Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, Justin Cui, Jiaqi Feng, Haoyu Xie, Tao Huang, Pichao Wang, Yanchao Yang, Cho-Jui Hsieh

arXiv 2609.28064首次发表:更新:

发表机构

The University of Sydney; University of California, Los Angeles; Shanghai Jiao Tong University; NVIDIA; The University of Hong Kong(悉尼大学; 加利福尼亚大学洛杉矶分校; 上海交通大学; 英伟达; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SlackDrive通过重用已实现延迟动态选择每个控制步骤的计算预算,在严格延迟约束下将NAVSIM v2上的EPDMS提升21.7%,有效缓解驾驶推理的实时性冲突。

AI 中文摘要

驾驶世界-动作模型通过将多模态推理与未来预测相结合来改进规划,但其不断增长的推理成本日益与车辆控制的实时延迟要求相冲突。现有的加速方法通过部署前选择的策略减少令牌、层或采样步骤,但在共享车载计算上进行离线分析和静态调度后,残留的运行时变化在很大程度上未被利用。我们观察到,最大允许计算预算随残留运行时状态系统性地变化,而近期实现的延迟提供了可用计算松弛的直接信号。基于这一观察,我们提出SlackDrive,一种推理前计算分配器,在模型执行前重用已实现的延迟来选择每个控制步骤的计算预算。SlackDrive对小型离散预算集的延迟和规划效用进行一次分析,从完成的向前传播中估计在线计算状态,并选择预测保持在允许延迟包络内的最高效用预算,补充现有的分析和资源调度,同时保留驾驶主干及其计算执行器。在NAVSIM v2上使用DriveDreamer-Policy,SlackDrive在严格延迟机制下将延迟约束的EPDMS比最强基线提高了21.7%,而全预算模型和预配置的令牌剪枝基线在运行时争用下超过了允许的延迟包络。

英文摘要

Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control. Existing acceleration methods reduce tokens, layers, or sampling steps with policies selected prior to deployment, yet leave residual runtime variation largely unexploited after offline profiling and static scheduling on shared onboard compute. We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack. Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution. SlackDrive profiles the latency and planning utility of a small discrete budget set once, estimates online compute state from completed forwards, and selects the highest-utility budget predicted to remain within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator. On NAVSIM v2 with DriveDreamer-Policy, SlackDrive improves latency-constrained EPDMS by $21.7\%$ over the strongest baseline under a stringent latency regime, while the full-budget model and preconfigured token-pruning baselines exceed the admissible latency envelope under runtime contention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑