arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

证据优先,算术其次:DocSem 系统报告与失败分析

Evidence First, Arithmetic Second: A System Report and Failure Analysis for DocSem

Divya Godara, Sachin Gupta

arXiv 2609.39013首次发表:更新:

AI 中文总结

本文报告 DocSem 系统 EVICALC 在官方测试中联合准确率仅 8.61%,通过人工失败分析发现 OCR 分块错误导致证据错位,并指出需进一步评估以确定原因。

AI 中文摘要

我们的 DocSem 共享任务系统 EVICALC 在官方最终测试评估中,于 1,730 个任务上取得了 8.61% 的联合准确率。该系统读取 PDF,选择一段文本,让语言模型编写一个算术表达式,并在本地代码中计算该表达式。保存的中间结果支持对失败进行检视。另一次公开验证运行取得了 92.17% 的答案准确率和 1.00 的证据 F1 分数。由于配置和指标不同,这些分数并非受控比较。我们的人工事后分析是描述性的:在一个被检视的案例中,光学字符识别(OCR)和分块分组将相关段落合并到了另一个块中,系统便根据无关文本作答。一项针对 100 份文档阅读页面图像的探索性研究仅为 22 份文档返回了证据标识符。这些描述性发现推动了进一步的评估;它们并未确定总体分数低下的成因。

英文摘要

EVICALC, our system for the DocSem shared task, achieved 8.61% joint accuracy on 1,730 tasks in the official final test evaluation. It reads a PDF, selects a passage, asks a language model to write an arithmetic expression, and evaluates that expression in local code. Saved intermediate results support inspection of failures. A separate public-validation run achieved 92.17% answer accuracy and 1.00 evidence F1. The configurations and metrics differ, so these scores are not a controlled comparison. Our manual, post-hoc analysis is descriptive: in one inspected case, optical character recognition (OCR) and block grouping merged the relevant passage into another block, and the system answered from unrelated text. An exploratory study of reading page images on 100 documents returned evidence identifiers for only 22 documents. These descriptive findings motivate further evaluation; they do not establish the causes of the overall score.

Comments5 pages, 1 figure, 2 tables. Accepted as a shared-task system paper at DocInsights 2026, co-located with EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑