相遇、比较或弃权:LatWeave 用于知识格上的确定性多跳问答
Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices
浏览论文内容
中文总结 AI 辅助
LatWeave 将知识组织为多维格,通过相遇、比较和弃权三个确定性算子实现多跳问答,确保答案路径零 LLM、可审计,并在完整知识下保持高性能,在信息不完整时可靠弃权。
中文摘要 AI 辅助
概率性问答系统——无论是大型语言模型(LLM)本身、检索增强生成(RAG),还是训练过的多跳检索器——都将“已知内容”和“如何推理”合并为单一的概率计算:幻觉无法根除,证据链无法审计,系统即使不知道答案也会回答。我们提出 LatWeave,它将知识组织成多维知识格,并将多跳问答编译为三个确定性算子——相遇(约束交集)、比较(格序比较)和弃权(结构性弃权);LLM 仅出现在构建侧(一次性提取)和查询规划侧,而答案生成路径是零 LLM、零任务训练且端到端可审计的——因此,对 Web 发布知识的问答变得逐项可复现。我们并未声称全面达到最先进水平,而是在六个公开基准上刻画了这一范式的运行范围:当知识完整时(MetaQA,39,093 个问题),相遇链在三跳内接近无损(any-hit 0.9975,与全监督 KBQA 相当);在模板化多跳主场上(2WikiMultihopQA 留出集 n=1,258),EM 为 0.865,远高于已发表的结构增强 RAG 复现结果;在开放文本深度组合(MuSiQue)和提取覆盖缺口(HotpotQA)上,我们如实报告了性能下降,并将其归因于格代数层之外的原因;当信息不完整时(IIRC),我们实现了结构性弃权,弃权准确率为 0.971,泄漏率为 0.029。在运行范围内,确定性执行不带来性能损失,且答案路径上的每一步都可重新计算——这正是端到端可审计性的来源。
英文摘要
Probabilistic question-answering systems -- whether large language models (LLMs) themselves, retrieval-augmented generation (RAG), or trained multi-hop retrievers -- conflate "what is known" and "how to reason" into a single probabilistic computation: hallucination cannot be eradicated, evidence chains cannot be audited, and the system answers even when it does not know. We present LatWeave, which organizes knowledge into a multidimensional knowledge lattice and compiles multi-hop QA into three deterministic operators -- meet (constraint intersection), compare (lattice-order comparison), and abstain (structural abstention); LLMs appear only on the construction side (one-shot extraction) and the query-planning side, while the answer-generation path is zero-LLM, zero-task-training, and auditable end to end -- so that question answering over Web-published knowledge becomes reproducible item by item. Rather than claiming across-the-board SOTA, we characterize the operating envelope of this paradigm on six public benchmarks: when knowledge is complete (MetaQA, 39,093 questions) meet chains are near-lossless over three hops (any-hit 0.9975, on par with fully supervised KBQA); on templated multi-hop home ground (2WikiMultihopQA held-out n=1,258) EM 0.865, well above published structure-augmented RAG reproductions; on open-text deep composition (MuSiQue) and extraction-coverage gaps (HotpotQA) we report degradation honestly and attribute it to causes outside the lattice-algebra layer; and when information is incomplete (IIRC) we achieve structural abstention with abstain accuracy 0.971 and leak rate 0.029. Within the operating envelope, deterministic execution pays no performance penalty, and every step on the answer path can be recomputed -- precisely the source of end-to-end auditability.
发表机构
- ZenSmart Technology (Beijing) Co., Ltd.(中数睿智科技(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。