发表机构
Tokyo Metropolitan University; Heinrich-Heine-Universität Düsseldorf(东京都立大学; 杜塞尔多夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RPyForth,通过向元追踪JIT编译器暴露固定宽度窗口和共享溢出区,实现Forth数据栈的有效编译,无需静态栈效应分析,显著提升性能。
AI 中文摘要
Forth是一种拼接语言,其词汇在调用之间共享一个数据栈,栈深度和调用效果在执行前无需知晓。这导致元追踪JIT编译器无法命名任何栈位置:一个单元通过栈指针访问,因此其访问保留在编译代码中,而声明栈数组的每个单元则会将编译器携带的信息与数组预留的容量绑定,而非与程序使用的深度绑定。我们的解决方案是:在栈顶使用一个固定宽度的窗口,一个共享的溢出区保存所有更深的单元,以及一个解码函数将两者连接起来,在此基础上,调用入口规范化和自适应入口是策略。解码表明,所有这些都保持了逻辑栈,并且追踪退出重建数据栈状态的范围受窗口宽度限制,而非栈深度。RPyForth在一个覆盖Forth核心词汇集的RPython解释器中实现了这一点,使用两个标量字段和八个帧位置,无需静态栈效应分析,而RPyFactor为Factor的子集实现了相同的窗口。将窗口暴露给编译器,而不仅仅是在其中缓存单元,才是关键所在。在布局和调用策略固定的情况下,对窗口字段进行注解使得十八个Shootout内核在两个x86-64机器上速度提升1.44-1.45倍,六个Appbench应用程序速度提升约1.60倍,RPyFactor中提升1.42-1.56倍。窗口的形状以及调用是否对其进行规范化影响较小,因程序而异,没有一种设置能在所有情况下获胜。作为一个完整的系统,RPyForth在两个套件上都比gforth-fast和SwiftForth更快,在内核上达到VFX Forth速度的1.90-2.36倍,而应用程序(其栈更深且调用更频繁)仍然是其弱点,速度为0.68-0.76倍。所有计时均针对已加载程序的重复执行。
英文摘要
Forth is a concatenative language whose words share one data stack across calls, with a depth and call effects that need not be known before execution. That leaves a meta-tracing JIT compiler with no stack location it can name: a cell is reached through the stack pointer, so its accesses stay in the compiled code, and declaring every cell of the stack array instead ties what the compiler carries to the capacity the array reserves rather than to the depth a program uses. We answer with a fixed-width window over the top of the stack, a shared spill holding every deeper cell, and a decoding function that joins the two, on top of which call-entry normalization and adaptive entry are policies. Decoding shows that all of them preserve the logical stack and that a trace exit rebuilds data-stack state bounded by the window's width, not by the stack's depth. RPyForth realizes this in an RPython interpreter covering Forth's Core word set, with two scalar fields and eight frame positions and no static stack-effect analysis, and RPyFactor realizes the same window for a subset of Factor. Exposing the window to the compiler, rather than merely caching cells in it, is what pays. With the layout and the call policy held fixed, annotating the window's fields makes eighteen Shootout kernels 1.44-1.45x faster and six Appbench applications about 1.60x faster on two x86-64 machines, and 1.42-1.56x in RPyFactor. How the window is shaped and whether calls normalize it matter much less, varying by program with no setting winning everywhere. As a complete system, RPyForth is faster than gforth-fast and SwiftForth on both suites and reaches 1.90-2.36x the speed of VFX Forth on the kernels, while the applications, whose stacks are deeper and whose calls are far more frequent, remain its weak point at 0.68-0.76x. All timings measure repeated execution of an already-loaded program.
Comments33 pages, 15 figures