发表机构
Tongji University; Swinburne University of Technology; Wuhan University of Technology(同济大学; 斯威本科技大学; 武汉理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建异构知识下LLM事实一致性基准TKFQA,提出训练框架ORLF,可提升LLM的推理与顺序鲁棒性,优于多种基线方法。
AI 中文摘要
大语言模型(LLMs)越来越多地支持基于用户提供的异构结构知识生成响应。然而,现有基准测试对LLMs能否在这类知识语境中忠实地执行多跳推理链,同时对输入顺序的变化保持鲁棒性的评估有限。我们推出TKFQA,这是一个包含10130个问答(QA)对的事实一致性基准,这些问答对基于表格、文本和知识图谱(KGs)构建。每个示例都由明确的反事实推理链构建而成,能够联合评估答案正确性、推理链准确性以及对不同输入顺序的鲁棒性。对14种开源和闭源LLMs的广泛评估显示,最先进的模型表现出有限的推理链准确性,并且对异构知识语境的输入顺序变化仍然敏感。为解决这些局限,我们提出ORLF,这是一种与LLM无关的训练框架,它通过特定于知识的潜在向量对跨上下文拓扑关系进行建模。ORLF整合了上下文位置编码、潜在桥接注意力掩码以及拓扑知识偏差,以保留特定于知识的偏差并编码拓扑语义。在四个LLM主干上的实验表明,ORLF优于有竞争力的无训练和基于LoRA的基线,使平均精确匹配值和推理链准确性分别提升2.15%和4.29%,同时将顺序诱导的性能标准差降低0.04%至3.01%。
英文摘要
Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. However, existing benchmarks provide limited assessment of whether LLMs can faithfully perform multi-hop reasoning chains across such knowledge contexts while remaining robust to variations in their input order. We introduce TKFQA, a factuality consistency benchmark comprising 10,130 question-answering (QA) pairs grounded in tables, texts, and knowledge graphs (KGs). Each example is constructed from an explicit counterfactual reasoning chain, enabling the joint evaluation of answer correctness, reasoning-chain accuracy, and robustness to different input-order. An extensive evaluation of 14 open- and closed-source LLMs reveals that state-of-the-art models exhibit limited reasoning-chain accuracy and remain sensitive to variations in the input order of heterogeneous knowledge contexts. To address these limitations, we propose ORLF, an LLM-agnostic training framework that models cross-context topological relations through knowledge-specific latent vectors. ORLF integrates context-wise position encoding, a latent-bridge attention mask, and topological knowledge bias to preserve knowledge-specific bias and encode topological semantics. Experiments across four LLM backbones show that ORLF outperforms competitive training-free and LoRA-based baselines, improving average Exact Match and Reasoning-Chain Accuracy by 2.15% and 4.29%, respectively, while reducing order-induced performance standard deviation by 0.04% to 3.01%.