发表机构
University of Toronto; Ontario Tech University(多伦多大学; 安大略理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对类型学特征预测现有方法可解释性不足且表现未充分探索的问题,采用基于URIEL+和Glottolog数据的上下文学习方法,发现结合系统发育与地理邻接证据的LLMs表现优于基线,且能生成符合证据的可解释依据。
AI 中文摘要
类型学特征在多语言自然语言处理中被广泛使用,预测这类特征对下游任务具有实用价值。然而,现有的缺失值预测方法缺乏对预测结果的可解释性依据,且其在不同资源水平和特征类型上的表现仍未得到充分探索。鉴于大型语言模型(LLMs)具备元语言推理和提供依据的能力,我们利用URIEL+和Glottolog的语言数据,通过上下文学习方法研究LLMs在类型学特征预测中的表现。我们发现零样本提示方式效果不足,但当提供系统发育和地理邻接证据时,LLMs的表现显著优于所有基线方法,且不会对低资源语言造成不利影响。我们还发现,大多数LLMs生成的依据与所提供的证据一致,为可解释的类型学特征预测迈出了一步。
英文摘要
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs' abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs' performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.
CommentsAccepted to EMNLP 2026