发表机构
University of Eastern Finland; Indian Institute of Technology Ropar(东芬兰大学; 印度理工学院罗帕尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于网络的框架表示八种语言的习语和比喻意义,用概念特征注释构建加权图,通过社区检测发现习语按概念模式聚类,网络能捕获独特语义信息,跨语言转移实验有提升,消融研究表明多特征维度有贡献,提供了可解释的习语意义表示。
AI 中文摘要
我们提出了一个基于网络的可解释框架,用于表示八种类型多样的语言中的习语和比喻意义,涵盖160个惯用表达式,其中大部分是习语。每个表达式都用从认知语言理论中得出的二元概念特征(包含、隐藏、情感、社会等)进行注释,成对的杰卡德相似度定义了一个加权图。社区检测表明,习语按概念模式而非语言聚类,产生了与认知语言预测一致的结构。概念网络捕获了分布嵌入中不存在的独特语义信息,可通过大语言模型自动注释进行扩展,改进了下游习语检测,并且在丰富语料库频率时保持稳健。跨语言转移实验表明,仅概念接近度就能识别五个语系中的可接受翻译对等物,比基于嵌入的基线有显著提升。消融研究表明,模式、角色和价这三个特征维度对网络的组织属性和习语检测性能都有非冗余贡献,并且特定的图衍生信号(社区成员身份、邻居相似度)特别有用。该框架提供了一种可解释的、跨语言稳定的习语意义表示,将理论基础与实际效用相结合。
英文摘要
We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic. Each expression is annotated with binary conceptual features (containment, concealment, emotional, social, etc.) derived from cognitive-linguistic theory, and pairwise Jaccard similarities define a weighted graph. Community detection reveals that idioms cluster by conceptual schema rather than by language, producing a structure consistent with cognitive-linguistic predictions. The conceptual network captures unique semantic information not present in distributional embeddings, can be scaled via automatic annotation with LLMs, improves downstream idiom detection, and remains robust when enriched with corpus frequencies. Cross-lingual transfer experiments show that conceptual proximity alone can identify acceptable translation equivalents across five language families, with substantial gains over embedding-based baselines. Ablation studies demonstrate that all three feature dimensions -- schemas, roles, and valence -- contribute non-redundantly to both the network's organizational properties and its performance on idiom detection, and that specific graph-derived signals (community membership, neighbor similarity) are particularly informative. The framework offers an interpretable, cross-linguistically stable representation of idiomatic meaning, combining theoretical grounding with practical utility.
CommentsSpelling of one of the author's name has been corrected