弃权错误与CUAD中的片段支持:一个可审计的合同抽取案例研究
Abstention Errors and Segment Support in CUAD: An Auditable Contract-Extraction Case Study
浏览论文内容
中文总结 AI 辅助
本文通过可复现的CUAD基准评估,揭示合同抽取系统在弃权(不执行)错误和类别支持度上的问题,强调审查负担需用候选级指标衡量。
中文摘要 AI 辅助
合同条款抽取基准衡量参考恢复能力,但评估用于审查辅助的系统还需要衡量不必要的输出以及类别内可用的证据。本文对固定的公共RoBERTa检查点在CUAD已发布的102份合同测试分割上进行了回顾性、可复现的评估。候选生成保留了先前冻结运行的结果;匹配错误被纠正,类别阈值在62份训练分割合同上使用原始的90%召回目标重新选择。纠正后的操作点在8.2%的CUAD精确率下恢复了82.2%的参考片段。在返回的30,464个候选字符串中,19,562个出现在没有标注答案的问题上。系统回答了2,938个此类问题中的1,335个,假阳性率为45.4%(95%合同自助法区间:43.3-47.9%)。匹配候选仅占返回字符串的20.3%,说明了为什么基于参考的精确率和候选级审查负担需要不同的分母。41个类别中只有9个拥有至少30个正例和30个负例合同;在支持度最小值20和40下,这一数量分别变为18和2。两个类别没有负例合同,使其无答案假阳性率未定义。这些发现描述了一个检查点、解码器和阈值策略。较早的测试访问和不确定的检查点训练重叠限制了确认性解释。本文提供了原始和纠正后的分析、预测及软件检查作为辅助工件;它既未确立新的发布标准,也未确立生产就绪性。
英文摘要
Contract-clause extraction benchmarks measure reference recovery, but interpreting a system for review assistance also requires measuring unnecessary output and the evidence available within categories. This paper presents a retrospective, reproducible evaluation of a fixed public RoBERTa checkpoint on CUAD's published 102-contract test split. Candidate generation is preserved from an earlier frozen run; matching errors are corrected and category thresholds are reselected on 62 training-split contracts using the original 90% recall target. The corrected operating point recovers 82.2% of reference spans at 8.2% CUAD precision. Of 30,464 returned candidate strings, 19,562 occur on questions with no annotated answer. The system answers 1,335 of 2,938 such questions, a 45.4% false-positive rate (95% contract-bootstrap interval: 43.3-47.9%). Matching candidates constitute 20.3% of returned strings, illustrating why reference-based precision and candidate-level review burden require different denominators. Only nine of 41 categories have at least 30 positive and 30 negative contracts; this count changes to 18 and two under support minima of 20 and 40. Two categories have no negative contracts, making their no-answer false-positive rates undefined. These findings describe one checkpoint, decoder, and threshold policy. Earlier test access and uncertain checkpoint training overlap limit confirmatory interpretation. The paper supplies the original and corrected analyses, predictions, and software checks as ancillary artifacts; it establishes neither a new release criterion nor production readiness.