arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38861cs.CLcs.DLcs.IR

TRACE:面向目标感知的检索、归因证据与契约约束抽取,用于LitTraceQA

TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA

Sachin Gupta, Divya Godara

AI总结:

TRACE通过目标分组检索、证据定位与契约约束抽取,解决LitTraceQA中源访问与评分可见正确性之间的差距,在官方测试集上取得0.760613的综合得分。

AI中文摘要:

找到相关论文并不等同于从中产生可验证的答案。LitTraceQA要求提供规范的论文标识符、页面或对象级别的精确证据,以及与评估器匹配的类型化答案。我们将源访问与评分者可见的正确性之间的分离称为“接地契约差距”。TRACE——即目标感知检索、归因证据与契约约束抽取——通过目标分组检索、独立类型化证据定位、多模态表格抽取、模式驱动的表格构建以及故障封闭验证来弥合这一差距。它通过段落、对象、别名、引文和稠密表示对27,487篇论文进行索引,同时保留每个信号背后的查询目标。对于表格,TRACE在提取数值之前预测观测单元,并使用与评估器兼容的键归一化来组装行。我们经过审计的选定干净轨道工件在官方71题测试集上得分为0.760613,包括0.9728的论文F1、0.6847的证据F1、0.9800的多项选择准确率、0.5423的表格行F1以及0.3508的宏单元格准确率。在11条公共开发表格记录上,干净基线和坐标感知视觉填充分别获得0.291和0.411的行F1;该诊断比较包含回退输出,并非官方测试声明。剩余错误主要涉及定位器、观测单元、行键和源值身份。

英文摘要:

Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract gap. TRACE - Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction - addresses this gap with target-grouped retrieval, independent typed evidence localization, multimodal table extraction, schema-driven table construction, and fail-closed validation. It indexes 27,487 papers through passage, object, alias, citation, and dense representations while retaining the question target behind each signal. For tables, TRACE predicts the observation unit before extracting values and assembles rows with evaluator-compatible key normalization. Our audited selected clean-track artifact scores 0.760613 on the official 71-question test set, including 0.9728 paper F1, 0.6847 evidence F1, 0.9800 multiple-choice accuracy, 0.5423 table-row F1, and 0.3508 macro cell accuracy. On 11 public-development table records, a clean baseline and coordinate-aware visual fill obtain row F1 of 0.291 and 0.411, respectively; this diagnostic comparison includes fallback outputs and is not an official-test claim. Remaining errors chiefly concern locator, observation-unit, row-key, and source-value identity.

补充信息

↑