发表机构
Jilin University; Peking University(吉林大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM在长篇叙事推理中易引入无支撑假说的问题,本文提出EVAR框架,通过编译证据库、分配推理预算及验证候选假说,在可控成本下提升了推理性能与证据忠实度。
AI 中文摘要
大型语言模型(LLM)在对非交互式长篇叙事进行推理时,常生成流畅但依据薄弱的结论。其核心失效模式是,无支撑的中间假说会进入推理轨迹并污染后续推理,尤其当证据分散在故事的不同远段时。为解决该问题,我们提出EVAR(一种面向预算感知叙事推理的证据验证假说准入框架)。EVAR首先将叙事编译为不可变的、关联源的原子声明证据库,并基于未解决的缺口和不确定性信号分配实例特定的推理预算。在细化过程中,EVAR针对未解决的缺口直接提出候选假说,构建依赖假说的验证挑战,并在准入前对照锁定的证据库验证每个候选:受支撑的假说进入答案支撑状态,无法验证的假说被隔离,矛盾的假说被丢弃。基于充分性的停止机制进一步避免不必要的细化。在NarraCrime及多个公共推理基准上的实验表明,EVAR在保持可控推理成本的同时,提升了任务性能与证据忠实度。
英文摘要
Large language models (LLMs) often produce fluent but weakly grounded conclusions when reasoning over non-interactive, long-form narratives. A central failure mode is that unsupported intermediate hypotheses can enter the reasoning trajectory and contaminate subsequent inference, especially when evidence is scattered across distant parts of the story. To address this problem, we propose EVAR, an evidence-validated hypothesis admission framework for budget-aware narrative reasoning. EVAR first compiles the narrative into an immutable evidence store of source-linked atomic claims and assigns an instance-specific inference budget from unresolved gaps and uncertainty signals. During refinement, EVAR directly proposes candidate hypotheses for unresolved gaps, constructs hypothesis-conditioned validation challenges, and verifies each candidate against the locked store before admission: supported hypotheses enter the answer-supporting state, unverifiable ones are quarantined, and contradictory ones are discarded. A sufficiency-based stopping mechanism further avoids unnecessary refinement. Experiments on NarraCrime and multiple public reasoning benchmarks show that EVAR improves both task performance and evidence faithfulness while maintaining controllable inference cost.
CommentsAccepted to the Main Conference of EMNLP 2026. 16 pages, 3 figures