arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StanceEval 2026:第二届立场检测共享任务

StanceEval 2026: The Second Stance Detection Shared Task

Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif

arXiv 2610.03215首次发表:更新:

发表机构

KFUPM; SDAIA–KFUPM Joint Research Center for Artificial Intelligence; HUMAIN(法赫德国王石油矿产大学; 沙特数据与人工智能局-法赫德国王石油矿产大学人工智能联合研究中心; HUMAIN)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

StanceEval 2026 是第二届阿拉伯语社交媒体立场检测共享任务,设跨目标和跨领域两个赛道,吸引多团队参与,顶级系统显著超越基线,且未见目标性能意外更高。

AI 中文摘要

StanceEval 2026 是 StanceEval 共享任务系列中关于阿拉伯语社交媒体文本立场检测的第二届。立场检测旨在识别作者对给定主题的立场。给定一条推文和一个目标,参与系统必须判断作者的立场是支持(Favor)、反对(Against)还是无(None)。本届专注于两个不同评估赛道中的跨目标泛化:赛道 1 评估主题相关的跨目标迁移(在训练数据中的“女性赋权”相关主题上训练,测试“女性驾驶”),而赛道 2 评估对完全未见目标的跨领域迁移(“电动汽车”和“学期制”)。该共享任务吸引了来自 12 个国家的 80 个注册团队。在评估阶段,30 个独特团队提交了参赛作品,经过验证过滤后,21 个团队在赛道 1 中正式排名,13 个团队在赛道 2 中正式排名,20 个团队提交了系统描述论文。参与团队采用了多样化的方法,包括微调的预训练语言模型、基于提示和检索增强的大型语言模型(LLM)、微调的 LLM 以及混合级联。顶级系统在赛道 1 上取得了令人瞩目的 $F_{avg2}$ 分数 0.8994,在赛道 2 上为 0.9400,大幅超过了最强基线(分别为 0.7366 和 0.7475),其中 $F_{avg2}$ 表示对支持(Favor)和反对(Against)类别进行宏平均的 F1 分数。与直觉相反,未见目标上的性能高于相关目标,这种差异可能由极端目标极化、类别不平衡以及跨主题的方言或讽刺细微差别驱动。

英文摘要

StanceEval 2026 is the second edition of the StanceEval shared task series on stance detection in Arabic social media text. Stance detection aims to identify a writer's stance toward a given topic. Given a tweet and a target, participating systems must determine whether the writer's stance is Favor, Against, or None. This edition focuses on cross-target generalization across two distinct evaluation tracks: Track 1 evaluates thematically related cross-target transfer (testing on Women Driving, related to Women Empowerment from training data), while Track 2 evaluates cross-domain transfer to completely unseen targets (E-Cars and Trimester System). The shared task attracted 80 registered teams from 12 countries. During the evaluation phase, 30 unique teams submitted entries, with 21 teams officially ranked in Track 1 and 13 in Track 2 following validation filtering, and 20 teams submitting system-description papers. Participating teams employed diverse methodologies, including fine-tuned pretrained language models, prompt-based and retrieval-augmented large language models (LLMs), fine-tuned LLMs, and hybrid cascades. Top systems achieved impressive $F_{avg2}$ scores of 0.8994 on Track 1 and 0.9400 on Track 2, substantially outperforming the strongest baselines (0.7366 and 0.7475, respectively), where $F_{avg2}$ denotes the macro-averaged F1 score over the Favor and Against classes. Counterintuitively, performance on the unseen targets was higher than on the related target, a disparity could be driven by extreme target polarization, class imbalance, and dialectal or sarcastic nuance across topics.

Comments14 pages total (8 pages main paper + 6 pages appendix), 5 tables in the main paper, excluding the appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑