arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

窗口未包含的内容:在基于文档的不稳定性基准中审计出处

What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark

Seyed Mosayeb Alam

arXiv 2609.06147首次发表:更新:

发表机构

KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建Probity基准并审计文档基准中证据缺失缺陷,发现被标记项目更不稳定,但修复窗口未能降低不稳定性,表明证据缺失与不稳定性仅相关而非因果。

AI 中文摘要

让语言模型就同一文档回答同一问题二十次,它有时会给出两个不同的答案。我们构建了Probity基准,包含来自真实风险投资备案文件的60个任务和470个项目,以衡量这种情况发生的频率。随后,我们审计了自己的语料库,发现任何基于摘录构建的基准都可能携带的一个缺陷:模型所看到的文本窗口中缺失了项目证据。审计标记了36个项目,并区分了单一标记会混淆的两种失败:证据确实不在窗口中,以及答案必须从窗口提供的数字中计算得出。被标记的项目改变答案的频率要高得多,在427个干净项目上波动率为0.255,而干净项目为0.087,排除它们会使表观跨模型一致性降低约五分之一。在测试缺失证据是否解释不稳定性之前,我们登记了一个预测:重新切割每个窗口以包含其证据,不稳定性应降至设定阈值以下。结果失败了:修复使波动率变化了0.058,其置信区间包含零。我们报告该关联为相关性。几乎所有测量都位于不稳定性无法显现之处,这限制了为准确性构建的语料库能对稳定性说明什么。我们发布了语料库、全部112,800条原始响应以及审计,作为任何文档基准的可运行检查。

英文摘要

Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built benchmark can carry: items whose evidence is missing from the window of text the model is shown. The audit flags 36 items and separates two failures a single flag would conflate: evidence genuinely absent from the window and answers that must be computed from numbers the window does supply. Flagged items change their answers far more often, wobbling at 0.255 against 0.087 on the 427 clean items, and excluding them cuts apparent cross-model agreement by about a fifth. Before testing whether the missing evidence explains the instability, we registered a prediction: re-cut each window to hold its evidence, and instability should fall below a set threshold. It failed: the repair moved wobble by 0.058, with an interval containing zero. We report the association as correlational. Almost all measurements sit where instability cannot show, which bounds what a corpus built for accuracy can say about stability. We release the corpus, all 112,800 raw responses, and the audit as a runnable check for any document benchmark.

CommentsAccepted as a poster at the 11th Workshop on Financial Technology and NLP (FinNLP 2026), co-located with EMNLP 2026. Code and data: https://github.com/eikiyo/FoFinNLP

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑