斯拉夫语中的说服技巧语料库
A Corpus of Persuasion Techniques in Slavic Languages
浏览论文内容
中文总结 AI 辅助
本文聚焦斯拉夫语构建说服技巧语料库,涵盖三国语言、约7500个文本跨度及25种细粒度技巧。描述创建过程并统计,研究话题与技巧相关性,还用多种模型提供检测和分类的基线及基准结果。
中文摘要 AI 辅助
说服技巧是用于在广泛媒体中影响公众舆论的强大修辞手段。我们提出了一个专注于斯拉夫语的说服技巧新语料库。该语料库包含保加利亚语、波兰语和俄语的文档,在粗粒度文本跨度级别和细粒度句子级别标注了说服技巧。这些技巧来自25种细粒度说服技巧的分类法,分为六大类修辞说服策略。语料库包含来自222个文档的约7500个文本跨度,涵盖国家和国际层面热议话题。我们描述了语料库创建过程,提供详细统计数据,并研究话题与说服技巧之间的相关性。我们使用基于经典机器学习和生成式人工智能的模型为文本跨度级别和句子级别说服技巧的检测和分类提供基线和基准结果。
英文摘要
Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of persuasion techniques, focusing on Slavic languages. The corpus contains documents in Bulgarian, Polish, and Russian, annotated with persuasion techniques at the coarse-grained text-span level and fine-grained sentence level. The techniques are drawn from a taxonomy of 25 fine-grained persuasion techniques, grouped under six broad categories of rhetorical persuasion strategies. The corpus contains approximately 7500 text spans from 222 documents that cover topics hotly debated at the national and international levels. We describe the corpus creation process, provide detailed statistics, and examine correlations between topics and persuasion techniques. We use classic ML-based and generative AI-based models to provide baselines and benchmark results for the detection and classification of persuasion techniques at the text-span level and sentence level.
发表机构
- Institute of Computer Science, Polish Academy of Science(波兰科学院计算机科学研究所)
- Sofia University "St. Kliment Ohridski"(索菲亚大学圣克莱门特·奥赫里德斯基分校)
- University of Koblenz(科布伦茨大学)
- Visa Technology Europe(维萨欧洲技术公司)
- CodeNLP(代码自然语言处理公司)
- University of Padua(帕多瓦大学)
- University of Helsinki(赫尔辛基大学)
机构由 AI 辅助整理,请以论文原文为准。