arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02391cs.CLcs.AI

PolERo:研究罗马尼亚语中的政治回避

PolERo: Studying Political Evasion in Romanian

Gabriel Stefan, Sergiu Nisioi

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建罗马尼亚语政治回避数据集PolERo,评估多种分类方法并研究跨语言迁移,发现微调编码器具竞争力、跨语言迁移不对称,矛盾回避类别是主要挑战。

中文摘要 AI 辅助

政治回避指的是在回应问题时保留所要求信息的行为。近期自然语言处理(NLP)研究将政治回避视为一项分类任务,采用回应清晰度与细粒度回避策略的两级分类体系。现有关于回应清晰度和回避分类的研究仅局限于英语,该分类体系与模型行为能否跨语言、跨政治语境迁移仍不明确。本文推出PolERo数据集,包含从五位罗马尼亚总统官方 transcript 中提取的3574个人类标注问答对。我们在匹配条件下对两个数据集评估多种分类方法,包括TF-IDF基线、微调编码器模型、本文提出的滑动窗口编码器,以及零/少样本大语言模型(LLM)提示。我们通过联合双语训练和基于机器翻译的数据增强研究跨语言迁移。结果显示,微调编码器具备竞争力,跨语言迁移呈不对称性,且涉及语用线索的矛盾回避类别仍是所有模型系列面临的主要挑战。

英文摘要

Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion strategies. Existing work on response clarity and evasion classification is limited to English, leaving open whether the taxonomy and model behavior transfer across languages and political contexts. We introduce PolERo, a dataset of 3,574 human-annotated question-answer pairs extracted from official transcripts of five Romanian presidents. We evaluate multiple classification approaches on both datasets under matched conditions, including TF-IDF baselines, fine-tuned encoder models, a proposed sliding-window encoder, and zero/few-shot LLM prompting. We study cross-lingual transfer through joint bilingual training and machine-translation-based data augmentation. Our results indicate that fine-tuned encoders are competitive, cross-lingual transfer is asymmetric, and ambivalent evasion categories involving pragmatic cues remain the main challenge across all model families.

发表机构

  • University of Bucharest(布加勒斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑