arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19153cs.CLcs.CYcs.IRcs.LG

停止移除停用词:继承的预处理默认设置如何扭曲法律文本数据

Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

Gregory M. Dickinson

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过单词语义消融实验证明,在可解释的法律文本分类流水线中,移除停用词(无论通用或优化列表)均无法提升F1分数,且可能扭曲教义信号,建议保留停用词以保障测量有效性。

中文摘要 AI 辅助

实证法律学术研究日益将司法文本视为数据,其中许多研究仍依赖稀疏、可解释的流水线——TF-IDF特征和线性分类器——因为文本特征本身往往是研究对象,而不仅仅是预测的手段。然而,这些流水线继承了20世纪中期信息检索中的一系列预处理默认设置,这些设置从未针对分类准确性进行验证,其中最根深蒂固的是停用词移除。本研究引入了一种详尽的单词消融方法,直接针对下游目标衡量预处理步骤的效果,并将其应用于停用词移除这一最难改变的情况。通过将最高法院数据库标签与Caselaw Access Project意见文本匹配,研究考察了两个跨越F1余量的二分类任务:意识形态方向(不移除基线F1约0.68)和宪法与非宪法法律类型(约0.92),涉及7,668和7,001份意见。对于每个任务,分析近似了任何专家能构建的最佳停用词表,移除约18,500个候选词中的每一个并直接衡量效果。研究得出三个发现:常用通用停用词表在所有测试中均低于不移除基线;即使优化后的停用词表在统计上与不移除无显著差异;基于词级特征训练的元模型无法预测哪些移除有帮助,因此停用词表整理没有明确目标。该方法可推广到任何继承的预处理默认设置,其结果对可解释的法律文本数据提出了具体警示:一个默默重塑模型所见特征的步骤,可能扭曲此类研究旨在恢复的教义和意识形态信号。保留停用词是一个测量有效性问题。

英文摘要

Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines (TF-IDF features and linear classifiers) because the textual feature is often the object of study rather than a means to a prediction. Yet these pipelines inherit preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy. The most entrenched of these is stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step's effect directly against the downstream objective, and applies it to stopword removal. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, the study examines two binary tasks, ideological direction (no-removal baseline F1 about 0.68) and constitutional versus non-constitutional law type (about 0.92), across 7,668 and 7,001 opinions. For each task, the ablation removes each of roughly 18,500 candidate words in turn, and a task-specific stoplist is built from the resulting measurements. Generic stoplists in common use fall below the no-removal baseline on held-out opinions in all twelve tests. The task-specific stoplists move held-out F1 by +0.0023 (95% CI [-0.0124, +0.0170]) on ideology and by +0.0001 ([-0.0082, +0.0085]) on law type. Neither task shows a detectable benefit from removal, and a supplemental analysis finds that word-level statistics predict a word's removal effect poorly, because the words' true removal effects differ by less than the measurement can register. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data, where a step that reshapes which features a model sees can distort the doctrinal and ideological signal the research is meant to recover. The burden of proof sits with removal.

发表机构

  • The University of Nebraska College of Law(内布拉斯加大学法学院)

机构由 AI 辅助整理,请以论文原文为准。

↑