LexLattice:基于文档层次结构的神经细胞自动机多语言抽取式摘要
LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
浏览论文内容
中文总结 AI 辅助
LexLattice通过二维神经细胞自动机整合法律文档层次结构,以1.8M参数实现多语言抽取式摘要的最先进ROUGE,并跨语言近无损迁移。
中文摘要 AI 辅助
忠实性是法律文本摘要中的一个核心关注点,这促使了抽取式方法的发展,这些方法选择可追溯至其来源的逐字内容。此类方法通常孤立地对段落或其他结构单元进行排序,却很少关注整合分布在文档遥远部分之间、且共享显著性的证据。我们引入了LexLattice,一种抽取式摘要器,它将法律法案的层次结构具体化为一个二维语义格,并在选择之前通过掩蔽的二维神经细胞自动机对其进行整合。LexLattice在EUR-Lex-Sum的所有24种语言的多语言和跨语言设置中均达到了最先进的ROUGE分数,超越了具有数十亿参数的指令调优基线,尽管其所有可训练能力都集中在冻结的多语言编码器之上的一个1.8M参数的整合器中。仅在高资源语言上训练的整合器进一步迁移到未见语言,且保留率近乎无损(0.99),这表明该模型 operates on 语言无关的语义几何而非表面形式。我们的结果将文档结构上的显式整合定位为多语言法律摘要中规模的一种紧凑且可追溯的替代方案。
英文摘要
Faithfulness is a central concern in legal text summarization, which motivates extractive approaches that select verbatim content traceable to its source. Such methods typically rank paragraphs or other structural units in isolation, yet give little attention to consolidating evidence that is distributed across, and shares salience between, distant parts of a document. We introduce LexLattice, an extractive summarizer that reifies a legal act's hierarchy as a two-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection. LexLattice attains state-of-the-art ROUGE across all 24 languages of EUR-Lex-Sum in both multilingual and cross-lingual settings, surpassing instruction-tuned baselines with billions of parameters, despite concentrating all trainable capacity in a 1.8M parameter consolidator over a frozen multilingual encoder. A consolidator trained only on high-resource languages further transfers to unseen languages with near-lossless retention (0.99), indicating that the model operates on language-agnostic semantic geometry rather than surface form. Our results position explicit consolidation over document structure as a compact and traceable alternative to scale for multilingual legal summarization.
发表机构
- The University of Winnipeg(温尼伯大学)
机构由 AI 辅助整理,请以论文原文为准。