发表机构
PricewaterhouseCoopers, U.S.(普华永道(美国))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在无人工证据标注下,仅用判决标签后训练小语言模型恢复证据,比较仅标签训练与拒绝采样,两者均提升证据恢复,但优劣需更多验证。
AI 中文摘要
在许多评审工作流中,判决是唯一被保留的内容。其背后的段落并未被标记,因为这种标注的成本远高于记录判决本身。我们衡量了当一个小型语言模型仅基于判决进行后训练、且在任一阶段均无人工证据标签时,它能恢复多少证据。在ContractNLI上,人工证据片段被保留至评估阶段才使用。匹配记录的判决与同意这些片段并非同一回事:在六个系统中,这两个分数仅弱相关,且对系统的排序不同,因此当引用需要可审查时,准确率是一个糟糕的指南。仅基于原始判决的标签训练达到准确率0.896和片段F1分数0.564。拒绝采样(仅在生成的轨迹与记录匹配时保留该轨迹,然后通过自动源接地分数选择一个)达到0.797和0.556,而训练前为0.747和0.493。逐字引用率从0.597上升至仅标签训练下的0.729和拒绝采样下的0.701。一个种子在一个语料库上无法说明哪种方法更好,但两者都在无人标注的情况下改善了证据。
英文摘要
In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage. On ContractNLI the human evidence spans are held out until evaluation. Matching the recorded verdict and agreeing with those spans are not the same thing: across six systems the two scores are only weakly related and rank the systems differently, so accuracy is a poor guide when the citations have to be reviewable. Label-only training on the bare verdict reaches accuracy 0.896 and span F1 0.564. Rejection sampling, which keeps a generated trace only when its verdict matches the record and then picks one by an automatic source-grounding score, reaches 0.797 and 0.556, against 0.747 and 0.493 before training. Verbatim citation rises from 0.597 to 0.729 under label-only training and to 0.701 under rejection sampling. One seed on one corpus cannot say which method is better, but both improve the evidence without anyone annotating it.
Comments16 pages, 5 figures