arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TabuLM:面向低资源语言的形态感知表格预训练模型

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre

arXiv 2608.26923首次发表:更新:

发表机构

Barclays; Carnegie Mellon University(巴克莱银行; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对低资源语言卢旺达语,提出形态感知表格预训练模型TabuLM,在自研的卢旺达语表格问答基准TabQA-kin上,其精确匹配率显著优于基线模型。

AI 中文摘要

我们提出TabuLM,这是首个基于卢旺达语(Kinyarwanda)表格数据预训练的语言模型。卢旺达语是一种屈折形态丰富的班图语,在卢旺达有超过1200万人使用,但缺乏任何专用的表格表示学习资源。TabuLM扩展了KinyaBERT-large,这是一个双层形态变换器,通过添加行、列和单元格类型嵌入,以及学习到的表格结构注意力偏置来强化同行和同列注意力。预训练使用两个新目标:掩码单元格恢复(MCR),即掩码整个单元格并要求从行和列的上下文中重建;以及列类型预测(CTP),即从观测到的单元格值预测列语义类型。我们在来自NISR、RAB、REB和MoH开放数据门户的172个卢旺达政府表格(约35000个单元格)上进行预训练,并引入TabQA-kin,这是首个原生卢旺达语表格问答基准,包含31个表格中的526个问答对,涵盖四种问题类型。TabuLM在TabQA-kin上达到62.0%的精确匹配(EM),比KinyaBERT-large高出5.7个EM点,比所有多语言基线(mBERT为49.3%,XLM-R为50.0%)高出11.7至12.7个点。分析表明,结构表格嵌入对比较和查找问题最具决定性,而形态感知则提供了互补的增益。我们的代码、数据和预训练检查点均公开可用。

英文摘要

We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is a morphologically rich Bantu language spoken by over 12 million people in Rwanda, yet lacks any dedicated tabular representation learning resource. TabuLM extends KinyaBERT-large, a two-tier morphological transformer, with additive row, column, and cell-type embeddings and a learned table-structure attention bias that sharpens same-row and same-column attention. Pre-training uses two new objectives: Masked Cell Recovery (MCR), which masks entire cells and forces reconstruction from row and column context, and Column Type Prediction (CTP), which predicts column semantic types from observed cell values. We pre-train on 172 Rwandan government tables (~35,000 cells) from NISR, RAB, REB, and MoH open-data portals, and introduce TabQA-kin, the first native Kinyarwanda table question-answering benchmark comprising 526 QA pairs across 31 tables and four question types. TabuLM achieves 62.0% exact match on TabQA-kin, outperforming KinyaBERT-large by 5.7 EM points and all multilingual baselines (mBERT 49.3%, XLM-R 50.0%) by 11.7-12.7 points. Analysis shows that structural table embeddings are most decisive for comparison and lookup questions, while morphological awareness provides complementary gains. Our code, data, and pre-trained checkpoint are publicly available.

Comments18 pages, 4 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑