发表机构
BRAC University; United International University; North South University; Spectrum Software & Consulting Ltd.(BRAC大学; 联合国际大学; 北南大学; 斯佩克特软件咨询有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对时间演进的文档理解构建了TIDE基准,评估9种LLMs在版本解析任务上的表现,发现模型在检测错误版本时表现较差,倾向于选择参数化答案而非权威文本。
AI 中文摘要
诸如法律、税法和软件文档等演进的文档会随时间被修订、替换,有时甚至会被撤销,因此同一问题在不同日期会有不同的正确答案。与旧事实被简单覆盖的百科知识不同,修订本身是官方文本,明确说明其替换内容和生效时间,而早期版本在其有效期内仍然有效。因此,核心挑战在于版本解析,即确定查询日期生效的版本。现有的时间问答数据集仅将时间作为注释,导致版本解析未得到测试。我们提出TIDE,这是一个经专家验证的基准,包含3050个问答对,涉及孟加拉国政府1969年至2025年间发布的644份官方海关文书,涵盖8种任务类型,涉及深度代码混合的文档,这些文档在布局上具有异质性,并采用两种历法标注日期。此外,我们在单一协议下评估了9种近期大语言模型(LLMs),分别在参数化、黄金上下文和检索访问三种设置下进行,由一个三人评审的LLM委员会评分,通过严格的日期门控区分正确含义和正确时间。最佳宏平均准确率仅为68.5%;从隐式日期解析版本的准确率为59.7%;检测提供的版本不管辖查询的准确率仅为26.7%。模型更有可能找到正确版本,而非拒绝错误版本,且它们倾向于选择自信的参数化答案,而非提供的权威文本。所有代码和数据可在此https URL获取。
英文摘要
Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates. In contrast to encyclopedic knowledge, where an old fact is simply overwritten, an amendment is itself an official text that states what it replaces and when it takes effect, and the earlier version stays correct for its validity period. The central challenge is therefore version resolution, that is, identifying the version in force on the queried date. Existing temporal QA datasets treat time only as an annotation, so version resolution stays untested. We present TIDE, an expert-verified benchmark of 3,050 QA pairs over 644 official customs instruments issued between 1969 and 2025 by the Government of Bangladesh, covering eight task types over deeply code-mixed documents that are heterogeneous in layout and dated in two calendars. In addition, we evaluate nine recent LLMs under a single protocol across parametric, gold-context, and retrieval access, scored by a three-judge LLM council with a hard date gate separating correct meaning from correct time. The best macro-averaged accuracy is only 68.5%. Resolving a version from an implicit date reaches 59.7%, and detecting that the supplied version does not govern the query reaches only 26.7%. Models are more likely to find correct versions than to reject incorrect ones, and they tend to follow a confident parametric answer over the supplied authoritative text. All code and data are available at https://github.com/icsetepa44/TIDE