发表机构
Autonomous University of Madrid (UAM); IE University; Oviedo University(马德里自治大学; IE大学; 奥维耶多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对阿拉伯语语言模型语义模糊问题,提出CAMMAR框架,通过分阶段语义课程组织意义到嵌套子空间,基于词汇与隐喻表示距离产生隐喻性几何度量,在新数据集上评估,有监督时检测隐喻效果好,还发现基于形态学词根有提升。
AI 中文摘要
阿拉伯语中的隐喻是一种基于文化构建意义、编码塑造解释的文化知识的机制。然而,当前阿拉伯语语言模型通常将词汇、文化和隐喻信息坍缩到单个表示空间,即“语义模糊”现象。我们引入了CAMMAR,这是一个表示学习框架,通过分阶段语义课程将意义组织到嵌套的词汇、文化和隐喻嵌入子空间中。该设计实现了阿尔 - 朱尔贾尼的nazum理论的组合原则,基于词汇和隐喻表示之间的距离产生了一种无需训练的隐喻性几何度量。在一个新的跨度注释阿拉伯隐喻集上进行评估,当层间几何结构由配对监督塑造时,几何读数能很好地检测隐喻(AUC高达0.84;82.6%的对中比喻得分高于其字面对应词),但仅在无监督域对比下处于随机水平,清晰区分了可解释的监督机制和无监督机制。控制消融表明,将词汇层基于形态学词根会带来小但一致的提升,直接探测中没有这种效果,这反映了该层作为测量锚的质量。我们将在接受后发布数据集、文化概念清单和代码。
英文摘要
Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet current Arabic language models typically collapse lexical, cultural, and metaphorical information into a single representational space, a phenomenon we term "semantic smearing". We introduce CAMMAR (Culture-Aware Matryoshka for Metaphorical Arabic Representations), a representation learning framework that organizes meaning into nested lexical, cultural, and metaphorical embedding subspaces through a staged semantic curriculum. The design implements compositional principles of Al-Jurjani's theory of nazum, modeling figurative meaning as compositionally grounded in prior semantic relations, and yields a training-free geometric measure of metaphoricity based on the distance between lexical and metaphorical representations. Evaluated on a new span-annotated Arabic metaphor set as word-matched figurative/literal pairs, the geometric readout detects metaphor well above chance when the inter-layer geometry is shaped by paired supervision (AUC up to 0.84; figurative outscores its literal counterpart for the same word in 82.6\% of pairs), but sits at chance under an unsupervised domain contrast alone, a clean separation between a legible-under-supervision regime and a non-emergent one. A controlled ablation shows that grounding the lexical layer in morphological roots gives a small but consistent gain, an effect absent from direct probing that reflects the layer's quality as a measurement anchor. We will release the datasets, cultural concept inventory, and code upon acceptance.
Comments14 pages, 5 figures