arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

乌尔都语中主要动词与轻动词区分的上下文嵌入证据

Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu

Farah Adeeba, Miriam Butt

arXiv 2608.23645首次发表:更新:

AI 中文总结

本研究借助UrduBERT等预训练模型的上下文嵌入,证实乌尔都语轻动词与其主要用法存在系统性差异但保留词干关联,且UrduBERT在轻动词动词身份预测任务中表现出良好的泛化能力。

AI 中文摘要

乌尔都语轻动词在保留与对应主要动词的词汇关联的同时,承担图式性事件结构意义。本研究使用UrduBERT、DunbaaBERT和multilingual BERT的上下文嵌入,对包含7个乌尔都语动词的1126个自然出现句子,测试源自Butt分析的表征预测。在所有21个动词-模型对比中,主要用法与轻动词用法表现出显著的表征分离;同时,同词干的主要与轻动词质心始终比不匹配的主要-轻动词词干对更接近,支持其持续的词汇关联。在仅针对轻动词用法的七分类预测任务中,目标被掩码后仍可恢复动词身份,UrduBERT达到0.866的准确率和0.852的宏F1值;在前置形式不重叠的评估下,UrduBERT仍保留0.782的准确率,表明其能泛化至重复局部动词组合之外的场景。这些发现提供了与Butt理论一致的计算证据,即乌尔都语轻动词与其主要用法存在系统性差异,同时保留词干特异性和动词特异性的表征结构。

英文摘要

Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. This study tests representational predictions derived from Butt's analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1,126 naturally occurring sentences containing seven Urdu verbs. Main and light uses show significant representational separation in all 21 verb--model comparisons. At the same time, same-lemma main and light centroids are consistently closer than mismatched main--light lemma pairs, supporting continued lexical relatedness. In a seven-way prediction task restricted to light uses, verb identity remains recoverable after the target is masked, with UrduBERT achieving 0.866 accuracy and 0.852 macro-F1. UrduBERT also retains 0.782 accuracy under a preceding-form-disjoint evaluation, indicating generalization beyond repeated local verb combinations. These findings provide computational evidence consistent with Butt's account that Urdu light verbs differ systematically from their main uses while retaining lemma-specific and verb-specific representational structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑