arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14617cs.LGcs.AIcs.CL

校准的信任,而非更精确的预测:不确定性融合的实证检验

Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

Surya Saka

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对1000个欧洲人权法院案件,在Claude Opus 4.8和GPT-5.5上测试含多种不确定性工具的融合流程,发现其未提升预测准确率但可通过共形选择性预测实现案件的自动化处理与升级,核心贡献是校准的信任而非更精确的预测。

中文摘要 AI 辅助

法律人工智能领域反复出现的一个提议是,将不确定性工具(带信念传播的证据图、序列贝叶斯优势更新、登普斯特-谢弗(Dempster-Shafer)组合及共形预测)融合为一条流程,以改进案件结果预测。我们针对LexGLUE和FairLex中的1000个欧洲人权法院真实案件,预测法院是否会根据案件事实段落判定存在《公约》违反行为。我们在两个前沿大语言模型(Claude Opus 4.8和GPT-5.5)上,将三个系列作为逐事实证据估计器进行比较:(A) 原始大语言模型,(B) 经融合流程路由的大语言模型,(C) 通过同一流程的词频基线。在约4750次测试中,我们发现:(1) 在歧视任务上(AUROC约为0.83),该流程未比原始大语言模型或基线产生任何改进;直接使用前沿大语言模型是最强的单一判别器。(2) 天真地将大语言模型与贝叶斯优势和登普斯特-谢弗融合组合,会因先验不匹配机制使校准误差增加一倍以上(预期校准误差(ECE)从约0.16升至0.46),该机制在两个模型中均存在。(3) 登普斯特-谢弗融合在长链上明显不安全,在准确率低于随机水平时仍会自信地给出错误标签;我们建议移除它。(4) 该流程的真正价值在于可操作性:通过共形选择性预测层路由,系统可决定哪些案件自动处理,哪些需升级处理。在移除登普斯特-谢弗、重新校准并对全部1000个案件集应用类条件风险控制后,调优后的引擎自动处理时准确率达96.8%,逃逸错误为0.5%,提交审查的案件占96.3%;相比之下,未调优基线的对应值为85.9%、3.8%、72.1%。此类流程在法律领域的贡献是校准的信任,而非更精确的预测。

英文摘要

A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the case's fact paragraphs. We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through the fusion pipeline, and (C) a term-frequency baseline through the same pipeline. Across roughly 4,750 tests we find: (1) on discrimination (AUROC around 0.83) the pipeline yields no improvement over either the raw LLM or the baseline; a frontier LLM used directly is the strongest single discriminator. (2) Naively composing an LLM with Bayesian-odds and Dempster-Shafer fusion more than doubles calibration error (ECE from about 0.16 to 0.46) via a prior-mismatch mechanism that replicates across both models. (3) Dempster-Shafer fusion is actively unsafe on long chains, committing confidently to wrong labels at below-chance accuracy; we recommend removing it. (4) The pipeline's genuine value is operational: routed through a conformal selective-prediction layer, the system decides which cases to automate and which to escalate. After removing Dempster-Shafer, recalibrating, and applying class-conditional risk control on the full 1,000-case set, the tuned engine auto-clears at 96.8 percent accuracy with 0.5 percent errors escaping and 96.3 percent caught for review, versus 85.9 / 3.8 / 72.1 for an untuned baseline. The contribution of such pipelines in law is calibrated trust, not sharper prediction.

发表机构

  • JudicialMind(司法智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑