arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文档提取的有效逐字段选择性风险控制:三种失效模式、一个有效性阶梯及条件化何时奏效

Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

Bhaskar Gurram

arXiv 2608.14639首次发表:更新:

发表机构

Zasti AI(扎斯蒂人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对文档提取的逐字段选择性风险控制问题,诊断出三种失效模式,提出有效性阶梯修复方案,通过多种方法验证了支持-bin等策略的效果,发布了开源代码。

AI 中文摘要

逐字段接受/审查的选择性风险最多为α——仅当接受字段中的错误率得到控制时才接受该字段——是文档提取系统所需的信任契约,但实际文档上的自然流程会悄然违反该契约。我们在来自800份CORD收据的13859个真实claude-sonnet-5字段(正确率为49.0%)上诊断出三种失效模式:文档聚类(设计效应为1.84-2.45)、分数再拟合泄漏(风险0.127时覆盖率为0.416,在95%的划分中违反α=0.10),以及平局质量病理(退化分数会崩溃阈值网格,从0.030降至0.001)。我们将修复措施组织为有效性阶梯,每个层级保证不同的形式。拟合/验证划分协议可恢复学习融合模型的预期选择性风险控制:名义α=0.10时,风险0.096下覆盖率为0.318,无容差带(生产变体为0.326)——这是一个平均点,其实现风险在47.5%的重划分中超过α,并非证书。采用精确二项式尾部的Mondrian Learn-then-Test可生成每组PAC证书:字段独立同分布(field-iid)时风险0.068下为0.171,聚类校正后为0.140,文档独立同分布(doc-iid)时为0.060——这是唯一与文档匹配的层级,目前几乎无实际应用。支持-bin(预指定的来源分类法)在sonnet CORD捕获的每个严格层级上均胜出(p<1e-4,经Bonferroni校正),但在同一文档上使用haiku或qwen时无法复现该优势;而在更高精度的语料库上,池化阈值胜出:条件化仅在池化无法认证的地方起作用,在其他地方则被学习分数取代。对未接触选择的claude-haiku-4-5进行冻结配置确认,并进行盲三标注员人工金标准审计,验证实际层级的接受集风险为1.3%,低于其10%预算(Fleiss' kappa=0.83;标签单侧偏向悲观)。发布Apache-2.0许可的代码,带有种子固定、回归门控流程。

英文摘要

Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real documents. On 13,859 genuine claude-sonnet-5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes: document clustering (design effect 1.84-2.45), score-refit leakage (coverage 0.416 at risk 0.127, violating alpha=0.10 in 95% of splits), and a tie-mass pathology (a degenerate score collapses the threshold grid, 0.030 to 0.001). We organize the fixes as a validity ladder, guarantee form stated per tier. A fit/val split protocol restores expected-selective-risk control for a learned fusion: coverage 0.318 at risk 0.096 at nominal alpha=0.10, no tolerance band (production variant 0.326) -- an on-average point whose realized risk exceeds alpha in 47.5% of resplits, not a certificate. Mondrian Learn-then-Test with exact binomial tails yields per-group PAC certificates: field-iid 0.171 at risk 0.068, cluster-corrected 0.140, doc-iid 0.060 -- the only tier matching documents, honestly near-vacuous today. Support-bin, the pre-specified provenance taxonomy, wins every rigor tier on the sonnet CORD capture (p<1e-4, Bonferroni-corrected) -- a win that does not replicate on the same documents under haiku or qwen -- while on higher-accuracy corpora pooled thresholds win: conditioning helps exactly where pooled cannot certify, subsumed by a learned score elsewhere. A frozen-configuration confirmation on selection-untouched claude-haiku-4-5 held at both risk levels, and a blind three-annotator human-gold audit verifies the practical tier's accepted-set risk at 1.3% against its 10% budget (Fleiss' kappa=0.83; labels err one-sidedly pessimistic). Released Apache-2.0 with seed-pinned, regression-gated procedures.

Comments14 pages. Seed-pinned, regression-gated harness (Apache-2.0): https://github.com/bhaskargurram-ai/verifydoc . Companion benchmark paper: VerifyDocBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑