混合自适应线程调优以缓解高性能强化学习推理中的仿真执行瓶颈
Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference
- College of Systems Engineering, National University of Defense Technology(国防科技大学系统工程学院)
- College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对仿真闭环RL推理的线程资源匹配问题,提出AutoThread方法,结合PINO与M/M/1排队模型并辅以在线微调,显著提升了推理的加速比与吞吐量。
AI中文摘要:
在仿真闭环决策系统中,强化学习(RL)推理常受仿真端执行开销限制,其工作负载高度动态且对运行时线程配置敏感。现有多线程策略难以在执行前或执行中匹配线程资源,导致资源竞争、调度开销增加及吞吐量降低。通过实证分析,本文确定任务执行时间与调度时间的比值是决定最优线程数的关键因素。基于该见解,本文提出AutoThread,一种用于缓解RL推理中仿真瓶颈的混合自适应线程调优方法。AutoThread采用物理信息神经算子(PINO)作为线程数预测器,并结合有限源M/M/1排队模型来约束和引导预测,以在动态工作负载下实现快速准确的估计;其还执行感知负载的在线微调,以补偿预测误差并优化资源分配。实验表明,AutoThread相比静态策略平均加速比提升18.4%,达到XGBoost和Reinforcer平均吞吐量的1.7倍和1.8倍,与最先进方法相比执行时间最多减少83.8%。本文的代码和数据集可在该https URL公开获取。
英文摘要:
In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling overhead, and reduced throughput. Through empirical analysis, we identify the ratio of task execution time to scheduling time as the key factor determining the optimal thread count. Building on this insight, we propose AutoThread, a hybrid adaptive thread-tuning method for mitigating simulation bottlenecks in RL inference. AutoThread employs a Physics-Informed Neural Operator (PINO) as a thread-count predictor and incorporates a finite-source M/M/1 queueing model to constrain and guide prediction, enabling fast and accurate estimation under dynamic workloads. It further performs load-aware online fine-tuning to compensate for prediction errors and refine resource allocation. Experiments show that AutoThread improves average speedup by 18.4\% over static strategies, achieves average throughput of 1.7x and 1.8x that of XGBoost and Reinforcer, respectively, and reduces execution time by up to 83.8\% compared with state-of-the-art methods. Our code and dataset are publicly available at https://github.com/suchenjm/AutoThread.