团结的智慧:多语言训练在谚语比喻语言识别中的作用
Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs
浏览论文内容
中文总结 AI 辅助
该研究以7种语言的6787个谚语实例为对象,评估多语言模型在比喻语言识别中的表现,发现约50%多语言训练数据可实现近最优性能,文化特异性比喻形式收益最大,推动相关框架向多维概念模型发展。
中文摘要 AI 辅助
尽管比喻语言识别的多语言方法并非新鲜事物,但超越语言同质训练数据的转变需要更清晰地理解翻译后的多语言监督的贡献。我们使用七种语言中6787个翻译实例所涵盖的742个谚语概念来研究这一问题。我们评估了五种模型,包括多语言编码器和指令调优的大型语言模型(LLM),在多语言监督水平逐步提升的情况下的表现。此外,我们引入了一种针对谚语的多维标注框架,通过四种互补的比喻形式对其进行表征:隐喻性、道德/建议性、因果性和文化特异性。我们的研究结果表明,约50%的翻译后多语言训练数据足以实现接近最优的比喻语言识别性能。我们进一步发现,结合不同的比喻形式可产生最强的整体性能。一个值得注意的发现是,出现频率最低的比喻形式——文化特异性形式,在多语言监督下表现出最大的性能提升。此外,道德/建议性和文化特异性形式对指令调优LLM在比喻语言识别任务中的性能贡献最大。这些发现推动多语言比喻语言识别超越以隐喻为中心的分类法,转向概念层面的多维框架,该框架明确建模互补的比喻意义形式。
英文摘要
Although multilingual approaches to figurative language identification are not new, the shift beyond language-homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We examine this question using 742 proverb concepts across 6,787 translated instances for seven languages. We evaluate five models including multilingual encoders and instruction-tuned LLMs through progressively increasing levels of multilingual supervision. Moreover, we introduce multidimensional annotation framework for proverbs that characterizes proverbs through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture-Specific. Our findings show that overall, adding multilingual training data beyond 50% provides only limited additional improvement, although the best supervision level varies across models and languages. Also, we show that combining diverse figurative forms yields the strongest overall performance. A notable finding is that the least frequent figurative form culture-specific exhibits the largest performance gains under multilingual supervision. Furthermore, the moral/advisory and culture-specific forms of proverb contribute more to instruct tuning LLM overall figurative identification performance. These findings motivate multilingual figurative identification to move beyond metaphor-centric taxonomies toward concept-level multidimensional frameworks that explicitly model complementary forms of figurative meanings that are context representative.
发表机构
- College of Computer and Information Sciences, King Saud University(沙特国王大学计算机与信息科学学院)
机构由 AI 辅助整理,请以论文原文为准。