arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03268cs.CL

Shrome at Touché:用于因果抽取的软投票集成与反因果增强

Shrome at Touché: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction

Roham Zendehdel Nobari, Shayan Sooratgar

首次发表
浏览论文内容

中文总结 AI 辅助

针对Touché 2026反因果新闻声明任务,提出每子任务单模型方案:检测用微调分类器加跨任务规则,抽取用三个RoBERTa-large标注器软投票集成,极性用大模型生成反因果句增强,在CCNC测试集上抽取F1达0.728居首。

中文摘要 AI 辅助

Touché 2026 将因果抽取任务扩展至反因果声明:即那些表面形式看似因果、但实际含义否定因果关系的新闻句子,例如“人们错误地认为 X 导致了 Y”。依赖诸如“caused”或“led to”等表面线索的系统会将此类句子误判为因果,并赋予其错误的极性。在反因果新闻语料库(CCNC)上,该任务包含三个子任务:判断句子是否为因果(检测)、定位其因果与结果跨度(抽取),以及将其极性标注为前因果、反因果或非因果。我们为每个子任务构建一个模型。检测采用一个微调的分类器,并应用一条跨任务规则,利用抽取出的跨度来去除误报。对于抽取,我们集成三个 RoBERTa-large BILOU+CRF 标注器,在解码前对其词元级分数进行平均,而非对各标注器生成的跨度进行投票。对于极性任务,由于标注的反因果示例最为稀缺,我们添加由大型语言模型生成的训练句子,该模型以改编自 Hagen 等人的九种反因果表达模式为提示,仅保留通过自动结构检查的句子。在留出的 CCNC 测试集上,系统在检测任务上达到 F1 0.869,在极性任务上达到宏平均 F1 0.817,并且在组织者最终仅限因果的抽取评估中,获得粒度调整后的 F1 0.728,是所有提交系统(包括组织者基线)中的最高抽取分数。开发集仅用于组件选择和论文中报告的消融实验。

英文摘要

Touché 2026 extends causality extraction to counter-causal claims: news sentences whose surface form appears causal but whose meaning denies the causation, as in "It is falsely believed that X caused Y." A system that relies on surface cues such as "caused" or "led to" will accept such a sentence as causal and give it the wrong polarity. On the Countercausal News Corpus (CCNC), the task has three subtasks: deciding whether a sentence is causal (detection), locating its cause and effect spans (extraction), and labeling its polarity as procausal, counter-causal, or uncausal. We build one model per subtask. Detection is a fine-tuned classifier with a single cross-task rule that uses the extracted spans to remove false positives. For extraction, we ensemble three RoBERTa-large BILOU+CRF taggers by averaging their token-level scores before decoding, rather than voting on the spans each tagger produces. For polarity, where labeled counter-causal examples are scarcest, we add training sentences generated by a large language model prompted with nine patterns of counter-causal expression adapted from Hagen et al., keeping only those that pass automatic structural checks. On the held-out CCNC test set, the system reaches F1 0.869 on detection and macro-F1 0.817 on polarity, and in the organizers' final causal-only evaluation of extraction it scores granularity-adjusted F1 0.728, the highest extraction score among all submissions including the organizers' baseline. The development split is used only for component selection and the ablations reported in the paper.

发表机构

  • University of Zurich(苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑