Padamitra:基于 grounded 术语表生成的古典梵语研究
Padamitra: Grounded Glossary Generation for Classical Sanskrit
AI总结:
该研究提出 grounded 术语表生成任务,构建含31316个三元组的梵语基准,发现指令微调及显式分词可提升性能,指出形态建模是梵词语法分解的关键瓶颈。
AI中文摘要:
我们提出了 grounded 术语表生成这一结构化任务,要求模型从颂歌-译文对中恢复具有语义意义的梵语短语,并生成与译文对齐的含义,将传统的 patha 评注实践形式化为可评估的 NLP 目标。我们从《罗摩衍那》和《薄伽梵往世书》中构建了包含 31316 个颂歌-译文-术语表三元组的基准,搭配两个指标:用于短语恢复的 Jaccard 指数和用于语义一致性的含义忠实度。针对 Gemma-3n-E4B、Gemma-3-12B、Phi-4 和 Qwen3.5-9B 的零样本、少样本及指令微调变体,指令微调的表现显著优于提示工程,而显式分词也带来了性能提升。错误分析表明,sandhi 和 samasa 复合词的过度分词是主要失败模式,这表明形态建模是实现忠实梵词语法分解的关键瓶颈。
英文摘要:
We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.