arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DPH解析器:一种用于联合成分分析和依存分析的由语法驱动的自底向上解析器

DPH Parser: A Bottom-Up Grammar-Driven Parser for Joint Constituency and Dependency Analysis

Hussein Ghaly

arXiv 2609.06070首次发表:更新:

AI 中文总结

本文提出DPH解析器,一种基于语法的自底向上无监督解析框架,通过特征规则和中心语标注联合生成成分与依存结构,在UD语料上验证了透明解析的可行性。

AI 中文摘要

本文提出了依存-短语层次解析器(DPH解析器),这是一种受广义短语结构语法(GPSG)和中心语驱动短语结构语法(HPSG)启发的、由语法驱动的自底向上无监督解析框架。该解析器使用紧凑的基于特征的句法规则库增量构建成分结构,同时通过显式中心语标注推导依存关系。该系统结合了概率词性标注、递归短语投影和加权解析假设,以处理现实且部分含噪的文本输入。与纯神经和数据驱动的解析器不同,其产生的句法推导过程保持显式可解释。我们使用通用依存(UD)项目中的英语语料库评估了解析器性能,以未标记依存准确率(UAS)作为主要解析指标,并将结果与Stanza和spaCy解析器进行了比较。对于一个小型句法规则库,DPH解析器在UD开发集和测试集上分别取得了53.32%和52.58%的UAS值。对于相同数据,Stanza分别取得了89.12%和88.67%,而spaCy分别取得了56.91%和58.59%。尽管当前系统尚未达到现代神经解析器的准确度,但结果证明了将透明的基于规则的自底向上解析应用于真实树库数据,并同时生成成分结构和依存结构的可行性。

英文摘要

This paper presents Dependency-Phrase Hierarchy Parser (DPH Parser), a grammar-driven bottom-up unsupervized parsing framework inspired by Generalized Phrase Structure Grammar (GPSG) and Head-driven Phrase Structure Grammar (HPSG). The parser incrementally constructs constituency structures using a compact inventory of feature-based syntactic rules while deriving dependency relations through explicit head annotations. The system combines probabilistic POS tagging, recursive phrase projection, and weighted parse hypotheses to process realistic and partially noisy text input. Unlike purely neural and data-driven parsers, the resulting syntactic derivations remain explicitly interpretable. We evaluated parser performance on English corpora from the Universal Dependencies (UD) project using Unlabeled Attachment Score (UAS) as the main parsing metric, comparing the outcomes against Stanza and spaCy parsers. For a small inventory of syntactic rules, DPH parser achieved UAS values of 53.32% & 52.58% (UD Devset/Testset respectively). For the same data, Stanza achieved 89.12% & 88.67% while spaCy achieved 56.91% and 58.59%. Although the current system does not yet approach the accuracy of modern neural parsers, the results demonstrate the feasibility of applying transparent rule-based bottom-up parsing to realistic treebank data while jointly producing constituency and dependency structures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑