arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MSLL:一种用于交互式文法开发的运行时多栈解析方法——LL风格递归下降的轻量级扩展

MSLL: A Runtime Multi-Stack Parsing Approach for Interactive Grammar Development - A Lightweight Extension of LL-Style Recursive Descent

Qunhui Zhang

arXiv 2609.19063首次发表:更新:

发表机构

School of Software, Shanghai Jiao Tong University(上海交通大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MSLL是一种轻量级运行时多栈解析方法,通过保留多个解析栈处理FIRST/FIRST冲突,支持交互式文法开发,实验表明在中小型输入上性能良好,作为LL解析器生成器的开发时辅助工具。

AI 中文摘要

很少有文法是按直线设计的。在文法原型设计过程中,作者通常希望在重构重叠的候选式或重新生成解析器代码之前,先运行一个部分确定的规则。预测性解析器对于稳定的文法非常有效,但其生成和编译循环可能会拖慢交互式文法开发。本文提出了MSLL,一种递归下降LL解析的轻量级运行时扩展,用于编辑时的文法探索。当出现FIRST/FIRST冲突时,MSLL保留多个活动的解析栈,让每个栈遵循不同的候选产生式,并在输入与某个栈矛盾时立即剪除该栈。因此,歧义被当作一个运行时状态管理问题来处理,而不是必须在执行前消除的条件。该原型面向ANTLR风格的文法工作流:直接运行并检查演化的文法,收集冲突轨迹,随后将稳定的文法翻译或适配到生成的解析器。在DSL风格工作负载上的实验表明,MSLL在48、49和59毫秒内完成了三个1k词法单元基准测试;在两个歧义的50k词法单元工作负载上保持在1秒以内;并在500k词法单元的深度嵌套压力测试中达到5.395秒。一个真实世界的JSON文法案例证实了在中小型输入上的交互行为,同时暴露了在500k词法单元时原型级别的超线性开销。结果将MSLL定位为现有基于LL的解析器生成器的开发时伴侣,而非生产解析器的替代品。

英文摘要

Few grammars are designed in a straight line. During grammar prototyping, authors often want to run a partially settled rule before refactoring overlapping alternatives or regenerating parser code. Predictive parsers are effective for stable grammars, but their generation and compilation loop can slow down interactive grammar development. This paper presents MSLL, a lightweight runtime extension of recursive-descent LL parsing for edit-time grammar exploration. When a FIRST/FIRST conflict appears, MSLL keeps multiple live parsing stacks, lets each stack follow a different candidate production, and prunes a stack as soon as the input contradicts it. Ambiguity is therefore handled as a runtime state-management problem rather than as a condition that must be removed before execution. The prototype targets an ANTLR-style grammar workflow: run and inspect an evolving grammar directly, collect conflict traces, and later translate or adapt the stabilized grammar for a generated parser. Experiments on DSL-style workloads show that MSLL finishes three 1k-token benchmarks in 48, 49, and 59 ms; stays under one second on two ambiguous 50k-token workloads; and reaches 5.395 s on a 500k-token deeply nested stress case. A real-world JSON grammar case confirms interactive behavior on small and medium inputs while exposing prototype-level super-linear overhead at 500k tokens. The results position MSLL as a development-time companion to existing LL-based parser generators rather than a replacement for production parsers.

Comments32 pages, 3 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑