arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合序列标注与方面代码切换增强的生成模型改进跨语言方面级情感分析

Generative Models Enhanced by Sequence Labelling and Aspect-Code Switching Improve Cross-lingual Aspect-Based Sentiment Analysis

Jakub Šmíd, Pavel Přibáň, Pavel Král

arXiv 2608.30425首次发表:更新:

发表机构

University of West Bohemia in Pilsen; NTIS – New Technologies for the Information Society; Faculty of Applied Sciences(比尔森西波希米亚大学; NTIS——信息社会新技术研究院; 应用科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SeqLab框架结合序列标注与方面代码切换技术,在11种语言、3个领域的跨语言ABSA任务中超越SOTA,还评估了多语言对并拓展至更具挑战的TASD任务。

AI 中文摘要

跨语言方面级情感分析(ABSA)将带有标注数据的源语言知识迁移至目标语言,无需目标语言标注数据即可实现细粒度情感分析。尽管单语ABSA已取得显著进展,但跨语言ABSA仍未得到充分探索,尤其在涉及目标-方面-情感检测(TASD)等包含多个情感元素的复杂任务中。本文提出一种新型SeqLab框架,该框架利用带辅助序列标注任务的序列到序列模型增强跨语言ABSA,其中编码器执行辅助序列标注任务以提升方面术语识别与情感预测能力。此外,本文引入方面代码切换(ACS)技术,这是一种基于翻译的方法,通过在源语言与翻译后的句子间交换方面术语生成额外训练数据,以增强模型的跨语言理解能力。本文在11种语言、3个领域及2种骨干模型上对所提方法进行评估,在常用的端到端方面级情感分析(E2E-ABSA)任务上超越了以往的最优结果。与多数仅依赖英语作为源语言的现有研究不同,本文系统评估了不同的源-目标语言对,并将评估扩展至跨语言环境中更具挑战性但未被充分探索的TASD任务。最后,本文提供了详细的错误分析,重点指出了关键挑战与局限性。

英文摘要

Cross-lingual aspect-based sentiment analysis (ABSA) transfers knowledge from a source language with annotated data to a target language, enabling fine-grained sentiment analysis without annotated target-language data. While monolingual ABSA has seen significant progress, cross-lingual ABSA remains underexplored, especially for complex tasks involving multiple sentiment elements like target-aspect-sentiment detection (TASD). In this paper, we propose a novel SeqLab framework that enhances cross-lingual ABSA using a sequence-to-sequence model with an auxiliary sequence-labelling task performed by the encoder, enhancing aspect term recognition and sentiment predictions. Additionally, we incorporate aspect-code switching (ACS), a translation-based technique that swaps aspect terms between source and translated sentences, generating additional training data to enhance the model's cross-lingual understanding. We evaluate our approach across eleven languages, three domains, and two backbone models, surpassing previous state-of-the-art results for the commonly studied E2E-ABSA task. Unlike most prior work that relies solely on English as the source language, we systematically assess different source-target language pairs and extend our evaluation to the more challenging, yet underexplored TASD task in cross-lingual settings. Finally, we provide a detailed error analysis highlighting key challenges and limitations.

CommentsAccepted for The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑