arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03473cs.FLcs.DCcs.PL

并行词法分析的已认证分割点:精确与丢弃标记模式

Certified Split Points for Parallel Lexing: Exact and Modulo Discarded Tokens

Nicklas Nidhögg

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出并行词法分析的已认证分割点条件,实现无需模拟等方法的并行扫描,在munch库测试中获92.6-95.3%并行效率与3.46-3.94倍端到端加速,将并行词法分析变为编译器可检查的属性。

中文摘要 AI 辅助

基于表驱动的DFA词法分析是串行的:每个转换依赖于前一字节的状态。要并行扫描一个输入,每个数据块的入口状态需被恢复,现有方法通过模拟、推测、预扫描或重叠来实现这一点。我们提出了无需上述任何方法的两个条件。对于在每个标记边界处从q0重启的最长匹配扫描器,当除q0外的所有可达状态都不存在可到达接受状态的b转换,且q0若存在此类转换则不可重入时,字节b即为已认证分割符号。完全可标记输入中此类字节的每一次出现都对应一个标记的起始,因此从此处开始的数据块通过有序拼接可复现串行序列的种类与长度。该条件既必要又充分,但具有脆弱性:一个字符串、注释或空白序列即可消除所有有用的认证,而注释和空白通常会被丢弃。因此,我们将保证弱化至删除声明的丢弃集合后的等价性,并提出第二个条件,该条件可靠且严格更具包容性但为保守而非精确,由相同的表判定,通过第二个常数时间的一位查询即可回答。它在不改变标记定义的情况下,为类C的常规词法分析恢复换行符,为JSON恢复制表符、换行符和回车符,在块注释不受限制的情况下则拒绝此操作。它仅作为查询提供:库的规划器及本文所有测量均使用精确条件,因此调用者必须自行规划边界。在munch库中,在受限CPU集上使用8线程,对超出最后一级缓存的512 MiB密集语料库,在精确认证处进行分割可达到92.6-95.3%的并行效率;在单台机器上的两个基准测试版本中,使用4线程可达到3.46-3.94倍的端到端加速。它将基于分隔符的并行词法分析从特定语言的假设转变为编译器可检查的属性。

英文摘要

Table-driven DFA lexing is sequential: each transition depends on the previous byte's state. Scanning one input in parallel needs each chunk's entry state, which existing methods recover by simulation, speculation, prescanning, or overlap. We give two conditions under which none is needed. For a longest-match scanner restarting from q0 at every token boundary, a byte b is a certified split symbol when no reachable state other than q0 has a b-transition whose target can reach acceptance, and q0 is not re-entrant if it has one. Every occurrence of such a byte in completely tokenizable input begins a token, so chunks starting there reproduce the serial sequence of kinds and lengths by ordered concatenation. The condition is necessary as well as sufficient, and fragile: one string, comment, or whitespace run can eliminate every useful certificate, and comments and whitespace are usually discarded. We therefore weaken the guarantee to equality after deleting a declared discarded set, and give a second condition, sound and more permissive, coinciding with the first when the discarded set is empty and strictly gaining on suitable pairs of token set and discarded set, but conservative rather than exact, decided from the same tables, answered by a second constant-time one-bit query. It recovers newline for a conventional C-like tokenization and tab, newline and carriage return for JSON, without altering their token definitions, and refuses it where block comments are unrestricted. Splitting at exact certificates in the munch library reaches 92.6-95.3% parallel efficiency at eight threads on a restricted CPU set, on a 512 MiB dense corpus beyond last-level cache, and a 3.46-3.94x end-to-end speedup at four threads, across two benchmark revisions on one machine. It turns delimiter-based parallel lexing from a language-specific assumption into a property a compiler checks.

补充信息

↑