AI 中文总结
该研究为确定有限自动机语言等价类配备超度量并识别相关空间,利用受保护语言算子的皮卡迭代收敛特性,构造深度受限自动机用于结构输入验证,并概述了实用的WAF管道,避免正则表达式和上下文无关解析器的问题。
AI 中文摘要
我们为确定性有限自动机的语言等价类配备了一个区分词超度量,并将所得空间等距地与正则语言进行识别。这个空间是不完备的,而它的度量完备化自然地与所有形式语言的完备超度量空间相识别。受保护语言算子在自动机空间上诱导收缩,其皮卡迭代在完备化中收敛到唯一的语言不动点,当它是正则的时候恰好由一个有限自动机表示。受结构输入验证的启发,我们使用这个框架来构造具有经认证的有限深度正确性的深度受限确定性有限自动机。这些自动机为嵌套输入结构(如带括号的SQL参数)提供了高效的预过滤器,同时避免了正则表达式引擎的回溯风险和全上下文无关解析器的运行时开销。我们还概述了一个结合了学习到的语法模型、有限状态构造和 \(O(1)\) 内存运行时验证的实用WAF管道。
英文摘要
We equip language-equivalence classes of deterministic finite automata with a distinguishing-word ultrametric and identify the resulting space isometrically with the regular languages. This space is incomplete, while its metric completion is naturally identified with the complete ultrametric space of all formal languages. Guarded language operators induce contractions on the automaton space, and their Picard iterates converge in the completion to the unique language fixed point, which is represented by a finite automaton exactly when it is regular. Motivated by structural input validation, we use this framework to construct depth-capped deterministic finite automata with certified finite-depth correctness. These automata provide efficient pre-filters for nested input structures, such as parenthesised SQL parameters, while avoiding the backtracking risks of regular-expression engines and the runtime overhead of full context-free parsers. We also outline a practical WAF pipeline combining learned grammar models, finite-state construction, and \(O(1)\)-memory runtime validation.