arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34263cs.AR

通过硬件/软件协同设计改进解释器中的间接分支预测

Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design

Linfeng Zheng, Hiroshi Sasaki

AI总结:

本文提出一种硬件/软件协同设计,利用硬件前瞻引擎和软件字节码元数据改善解释器的间接分支预测,仅需1.3 KB存储和少量代码修改,在CPython服务器负载上显著降低MPKI并提升IPC。

AI中文摘要:

解释器具有较大的间接分支足迹,需要较大的预测器容量才能进行准确预测。我们提出了一种硬件/软件协同设计,其中硬件前瞻引擎在流水线之前运行,并利用软件提供的字节码元数据,向前端提供解释器调度目标。该引擎仅需1.3 KB的片上存储,并需对约50行CPython代码进行修改。在15个CPython服务器工作负载上,一个14 KB的ITTAGE(增强型间接分支预测器)结合该引擎,相对于16 KB的ITTAGE基线,将字节码跳转的MPKI(每千条指令的未命中数)降低了73.7%,从而实现了3.2%的调和平均IPC加速。

英文摘要:

Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.

补充信息

↑