发表机构
Sri Sivasubramaniya Nadar College of Engineering; National Institute of Technology Tiruchirappalli(斯里·西瓦苏布拉马尼亚·纳达尔工程学院; 蒂鲁吉拉伯利国家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建古典泰米尔语诗句-评注对语料库,训练多种模型探究表示学习的信息恢复能力,发现仅解码器模型偏好真实词序但无法复现评注内容,发布了提取与评估协议。
AI 中文摘要
我们从5个古典泰米尔语文本片段中构建了包含1262对诗句-评注(urai)的语料库,涵盖从技术语法散文到现代意译的内容,探究表示学习能恢复何种信息。我们训练了循环编码器、Transformer编码器、连体式配对匹配网络、mBART式编码器-解码器以及仅解码器的语言模型,每项分析均对照同一数据上的合适对照组进行解读。TF-IDF提供了无需训练的强词汇检索基线,同时为学习到的模型提供表示分析和生成对照。包含25个最常见评注词的固定字符串在生成重叠度上得分高于仅解码器模型;在当前样本量下,高斯噪声的典型相关系数达到1.000;token-F1在该语料库上仅为0.02-0.20左右;编码器-解码器在验证损失已开始上升后,仍在16个epoch中持续降低训练损失。一项明确结果是:在112对最小配对比较中,仅解码器模型在107对(95.5%)中偏好真实词序,但无法复现保留的评注内容。我们发布了提取与评估协议;源评注的重新分发需遵守许可。
英文摘要
We construct a corpus of 1,262 verse--commentary (urai) pairs from five Classical Tamil source sections, ranging from technical grammatical prose to modern paraphrase, and ask what information representation learning can recover. We train recurrent and Transformer encoders, a Siamese-style pair-matching network, an mBART-style encoder--decoder, and a decoder-only language model. Each analysis is interpreted against an appropriate control on the same data. TF-IDF provides a strong no-training lexical retrieval baseline, alongside representation analyses and generation controls for the learned models. A fixed string containing the 25 most frequent commentary words scores higher on generation overlap than the decoder-only model. Canonical correlation reaches 1.000 on Gaussian noise at these sample sizes, token-F1 spans only about 0.02--0.20 on this corpus, and the encoder--decoder continues to lower training loss for sixteen epochs after validation loss has begun to rise. One narrow result remains: the decoder-only model prefers authentic word order in 107 of 112 minimal-pair comparisons (95.5%), but does not reproduce held-out commentary content. We release the extraction and evaluation protocol; redistribution of the source commentaries remains subject to permission.
Comments8 pages, 5 figures, 6 tables