arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用依存图和邻近特征从历史报纸中进行轻量级人物-地点关系提取

Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features

Mlen-Too Wesley

arXiv 2607.19718首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从历史报纸中提取人物-地点关系,构建文档级图,提取实体对特征,用小型模型分类。在官方评估中取得一定成绩,发现最小字符距离重要,且文档分组交叉验证对避免数据泄漏、保证结果可靠至关重要。

AI 中文摘要

HIPE-2026共享任务引入了从多语言历史报纸中提取人物-地点关系作为一个新的评估赛道,对英文、法文和德文中预先标注的人物和地点提及之间的“at”和“isAt”关系进行分类。受大规模处理历史档案成本的推动,我们团队(DS@GT HIPE,官方结果中的团队2)研究了在关系分类阶段,一个轻量级、可解释的系统在不使用任何预训练语言模型的情况下能走多远。我们的方法从依存解析构建文档级图,为每个实体对提取基于邻近性和词性的特征,并用小型的scikit-learn集成模型或紧凑的图注意力网络对其进行分类,每次提交运行的参数都保持在847K以下。在官方评估(测试A,报纸测试集)中,我们的最佳运行达到了0.5142的宏召回率,在效率方面排名第3,在17个参赛团队的准确性方面处于中游。有两个发现很突出。第一,仅最小字符距离就捕获了大部分分类信号;添加更多工程特征带来的收益不一致,有时还会降低性能,这与之前论证距离主导关系提取的证据相呼应。第二,文档分组交叉验证在这个语料库上至关重要:由于实体提及在文档中反复出现,成对级别的分割会使分数膨胀25 - 37个百分点,而分组交叉验证消除了这种数据泄漏效应。

英文摘要

The HIPE-2026 shared task introduces person-place relation extraction from multilingual historical newspapers as a new evaluation track, classifying the at and isAt relations between pre-annotated person and location mentions in English, French, and German. Motivated by the cost of processing historical archives at scale, our team (DS@GT HIPE, team 2 in the official results) investigates how far a lightweight, interpretable system can go without any pretrained language model at the relation classification stage. Our approach builds a document-level graph from dependency parses, extracts proximity-based and part-of-speech features for each entity pair, and classifies them with small scikit-learn ensembles or compact Graph Attention Networks, keeping every submitted run under 847K parameters. On the official evaluation (Test A, the newspaper test set), our best run reached a macro recall of 0.5142, ranking 3rd on the Efficiency profile while placing mid-table on Accuracy among the 17 participating teams. Two findings stand out. First, minimum character distance alone captures most of the classification signal; adding further engineered features yields inconsistent gains and sometimes degrades performance, echoing prior evidence that argument distance dominates relation extraction. Second, document-grouped cross-validation is essential on this corpus: pair-level splits inflate scores by 25-37 percentage points because entity mentions recur across documents, a data-leakage effect that grouped cross-validation removes.

Comments19 pages, 4 figures. Accepted at CLEF 2026 HIPE Shared Task. To appear in CEUR Workshop Proceedings (CEUR-WS.org)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑