发表机构
Indian Institute of Technology Kharagpur; Central Sanskrit University; Manipal Academy of Higher Education(印度理工学院卡拉格普尔分校; 中央梵语大学; 马尼帕尔高等教育学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出首个基于正理-胜论本体的梵文细粒度NER基准Padārtha,标注《摩诃婆罗多》等语料,发现微调生成式NER模型与特定任务系统性能相当,但从粗到细粒度时性能骤降且难处理未见过的实体。
AI 中文摘要
标注模式并非中立的。当将为现代新闻文本开发的标签集应用于古典文学时,它们会将源文化的定义强加于从未设计用于描述的文本之上。相反,我们将模式建立在文本自身的传统之上,推出了Padārtha——首个基于本体的古典梵文细粒度命名实体识别(NER)基准,构建于《摩诃婆罗多》(Mahābhārata)史诗之上。我们的标签集源自印度古典本体论体系正理-胜论(Nyāya-Vaiśesika),产生了18个细粒度类别,归属于10个本体节点,并映射到5个标准粗标签,确保与现有基准的互操作性。专业标注员对来自命名实体学术索引的超过12600个条目进行标注,这些条目与《摩诃那摩》(Mahānāma)语料库中的相应提及关联,生成了73632节经文中共108335个实体提及的细粒度标注,以及一个抽样以突出罕见提及的5000节经文的专家验证测试集。我们首次对梵文的生成式NER与传统架构进行了系统基准测试,发现微调后的生成式模型表现与特定任务系统相当。然而,所有系统从粗粒度到细粒度时均出现性能急剧下降,且难以处理训练期间未出现的实体提及。该局限并非仅因数据稀缺,因为微调模型对未见过实体的召回率远差于见过的实体,且在词汇歧义下倾向于默认采用多数含义。
英文摘要
Annotation schemas are not neutral. When applied to classical literature, tag sets developed for modern journalistic texts impose source-culture definitions on texts they were never designed to describe. We instead ground a schema in the tradition of the text itself introducing \textit{Padārtha}, the first ontology-grounded fine-grained Named Entity Recognition (NER) benchmark for Sanskrit, built on the \textit{Mahābhārata} epic. Our tag set derives from \textit{Nyāya-Vaiśesika}, a classical Indian ontological system, yielding 18 fine-grained categories organized under 10 ontological nodes and mapped onto five standard coarse tags, ensuring interoperability with existing benchmarks. Expert annotators label over 12.6K entries from a scholarly index of named entities, linked to corresponding mentions in the \textit{Mahānāma} corpus, producing fine-grained annotations for 108,335 entity mentions across 73,632 verses, along with a 5,000-verse expert-verified test set sampled to stress rare mentions. We present the first systematic benchmarking of generative NER against traditional architectures for Sanskrit, finding that fine-tuned generative models perform comparably to task-specific systems. However, all systems show a sharp decline from coarse to fine granularity and struggle with out-of-entity mentions unseen during training. The limitation is not due to data scarcity alone, as fine-tuned models recall unseen entities far worse than seen ones and tend to default to the majority sense under lexical ambiguity.