arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在黑盒设置中为大型语言推断准确的上下文无关语法

Toward Inferring Accurate Context-free Grammars for Big Languages in a Black-box Setting

Mohammad Rifat Arefin, Nuhiat Arefin, Shanto Rahman, Christoph Csallner

arXiv 2607.08959首次发表:更新:

AI 中文总结

针对黑盒上下文无关语法推断在大型语言中存在的可扩展性、准确性和语法可读性问题,Xvada引入新技术,在实证比较中语法准确性和紧凑性优于对手,还发现Python Liquid引擎漏洞及多个新漏洞,且相关内容可免费获取。

AI 中文摘要

黑盒上下文无关语法推断对程序分析、逆向工程、程序理解、模糊测试和安全至关重要。但现有方法如Arvada、TreeVada、Kedavra和Cucio在可扩展性、准确性和语法可读性方面存在问题,尤其针对大型语言。为应对挑战,Xvada引入新技术进行上下文无关语法的确定性推断。在实证比较中,Xvada在语法准确性和紧凑性上优于最高得分的竞争对手TreeVada。还发现Python Liquid引擎的CVE,基于Xvada推断语法的模糊测试又发现五个漏洞并被修复。XVADA及实验数据和脚本可免费获取。

英文摘要

Black-box context-free grammar inference is crucial for program analysis, reverse engineering, program understanding, fuzzing, and security. But existing approaches such as Arvada, TreeVada, Kedavra, and Cucio struggle with scalability, accuracy, and grammar readability, especially on larger languages. To address this challenge, we introduce XVada with several new techniques for deterministic inference of context-free grammars. In an empirical comparison that avoids several pitfalls of recent studies, XVada improves on the highest-scoring competitor (TreeVada) both in grammar accuracy and grammar compactness. XVada also found a CVE in the widely used Python Liquid engine. Fuzzing based on the XVada-inferred grammar found five more bugs, which the Python Liquid developers fixed based on our bug reports. XVada and all experimental data and scripts are freely available.

Comments12 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑