发表机构
Faculty of Law, University of Cambridge; The Psychometrics Centre, Cambridge Judge Business School, University of Cambridge; Faculty of English, University of Cambridge; Department of Statistics, Uppsala University; Department of Computer Science and Technology, Tsinghua University; Department of Engineering, University of Cambridge(剑桥大学法学院; 剑桥大学心理学测量中心,剑桥Judge商学院; 剑桥大学英语学院; 乌普萨拉大学统计系; 清华大学计算机科学与技术系; 剑桥大学工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究英国就业法庭判决中索赔结果预测的捷径学习,用33158条索赔语料库预测结果,发现基于事后司法文本训练的法律判决预测系统性能或被夸大,去除泄漏特征后模型仍能提取有用信号。
AI 中文摘要
当前法律判决预测依赖事后司法材料,易进行回顾性分类而非真正预测。本文通过研究英国就业法庭索赔级结果预测实证调查捷径学习。虽预测性能指标看似良好,但基于事后司法文本训练的系统性能可能受材料性质影响。分层测试数据发现泄漏线索影响性能,去除泄漏特征训练的模型表现良好。这表明性能可能被语言假象夸大,但并非致命,去除假象后模型仍能提取有用信号。
英文摘要
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.