arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UNMASK:发现并因果验证文本分类器中的虚假捷径

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Chidaksh Ravuru, Shashank Srivastava

arXiv 2608.09209首次发表:更新:

发表机构

University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出UNMASK自动化流程,可发现并因果验证文本分类器的虚假相关性,在MNLI、CivilComments-WILDS等数据集上提升模型性能,还可推广至奖励模型偏好数据。

AI 中文摘要

在大规模众包语料库上训练的神经语言模型,常利用与目标标签相关联的虚假表面模式,这些模式不具备真正的语言学或因果相关性,虽能提升基准性能,但在对抗性输入或分布外输入上表现糟糕。现有方法要么需要手动指定特征词汇,要么仅能部分实现自动化发现,未解决数据集层面相关性与模型层面利用之间的差距问题。本文提出UNMASK,这是一种完全自动化的流程,无需额外人工标注即可发现、因果验证并缓解文本分类器中的虚假相关性。给定未标注的训练样本,UNMASK会生成候选表面模式作为可执行布尔表达式,通过带独立重复的统计验证协议对其进行过滤,并通过验证的反事实干预建立因果模型依赖性。经因果确认的特征随后用作无标注的组定义,用于深度特征重加权(Deep Feature Reweighting),消除了标准深度特征重加权所需的组标签。将该流程应用于在MNLI上训练的BERT和RoBERTa,其独立重新发现了已确立的词汇重叠和否定偏差,验证了BERT上10个特征中的9个、RoBERTa上的6个,使HANS准确率提升了12.58个百分点。在CivilComments-WILDS上,该流程生成的程序化组匹配了带人工标注的深度特征重加权(Kirichenko等人,2023)的70.1%最差组准确率,且无需人口统计标注。本文进一步证明,该发现与验证阶段可推广至奖励模型偏好数据,在RewardBench2中揭示了可解释的虚假相关性。

英文摘要

Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.

CommentsAccepted at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑