WILDTRACE:长文本推理中自然证据线索的基准测试
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
浏览论文内容
中文总结 AI 辅助
该研究针对长文档复杂问题回答中源内证据整合的缺失,引入WILDTRACE基准,利用因果层次等定义证据几何结构,通过源优先管道挖掘线索并验证,为长文本研究应对信息与推理差距提供基准。
中文摘要 AI 辅助
回答长文档中的复杂问题通常需要整合分散在远处段落的自然证据。现有基准大多回避了这种源内证据整合问题。我们引入了WILDTRACE基准,它包含481个任务,涉及214个自然出现的长文本源。利用因果层次和多跳推理类型,定义了七种源内证据几何结构。通过源优先构建管道挖掘候选线索,并进行多阶段验证。随着模型承担更多现实世界的分析任务,获取信息与对自然分散证据进行推理之间的差距成为长文本研究下一阶段的关键挑战。
英文摘要
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character's true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic. Drawing on Pearl's causal hierarchy and prior multi-hop reasoning typologies, we define seven source-internal evidence geometries that characterize the distinct relational demands of analytical reading in long documents. A source-first construction pipeline mines candidate trails from document structure before writing questions; each item then undergoes multi-stage validation covering clue necessity, answer groundedness, rubric fidelity, contamination resistance and answerability. As models are increasingly entrusted with real-world high-stakes analytical tasks, this gap between accessing information and reasoning over naturally dispersed evidence emerges as a defining challenge for the next stage of long-context research.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Qwen Team, Alibaba Group(阿里集团通义团队)
机构由 AI 辅助整理,请以论文原文为准。