发表机构
National University of Defense Technology; Harbin Institute of Technology, Shenzhen(国防科技大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究开放域多跳问答中最终答案选择难题,提出STEC证据压缩框架从候选集选答案。通过答案级证据压缩和证据引导的答案验证机制,将选择从原始轨迹比较转为候选级证据比较,实验显示其性能最佳。
AI 中文摘要
在开放域多跳问答(QA)中,基于大语言模型的搜索代理通过结合检索与推理,为知识密集型QA提供了一种很有前景的方法。现有方法主要通过推理范式、检索交互和搜索策略优化来改进开放域多跳QA。然而,使用多条搜索轨迹会带来具有挑战性的最终答案选择问题。不同轨迹可能支持不同候选答案,检索到的信息可能是异构、冗余、不完整或冲突的。直接比较原始轨迹会使验证器暴露于噪声和未对齐的内容,而比较答案字符串则忽略了支持每个候选答案的证据,使得可靠的最终选择变得困难。为应对这一挑战,我们提出了STEC,一种用于多跳QA中最终答案选择的证据压缩框架。STEC通过两种机制从现有候选集中选择最终答案:(1)答案级证据压缩,按归一化答案标识对轨迹进行分组,并将每个答案组转换为特定于候选答案的证据表示;(2)证据引导的答案验证,比较这些表示并从候选集中选择最终答案。这种设计将最终选择从原始轨迹比较转移到候选级证据比较。我们在四个开放域多跳QA基准上针对代表性基线评估了STEC。实验结果表明,STEC在比较的方法中总体表现最佳,消融结果证明答案级证据压缩有助于最终答案选择。
英文摘要
In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach to knowledge-intensive QA by combining retrieval with reasoning. Existing methods mainly improve open-domain multi-hop QA through reasoning paradigms, retrieval interaction, and search strategy optimization. However, using multiple search trajectories introduces a challenging final answer selection problem. Different trajectories may support different candidates, and the retrieved information can be heterogeneous, redundant, incomplete, or conflicting. Directly comparing raw trajectories exposes the verifier to noisy and unaligned content, while comparing answer strings ignores the evidence supporting each candidate, making reliable final selection difficult. To address this challenge, we propose STEC, an evidence compression framework for final answer selection in multi-hop QA. STEC selects the final answer from the existing candidate set through two mechanisms: (1) Answer-Level Evidence Compression, which groups trajectories by normalized answer identity and converts each answer group into a candidate-specific evidence representation; and (2) Evidence-Guided Answer Verification, which compares these representations and selects the final answer from the candidate set. The design shifts final selection from raw trajectory comparison to candidate-level evidence comparison. We evaluate STEC on four open-domain multi-hop QA benchmarks against representative baselines. Experimental results show that STEC performs best overall among the compared methods, and ablation results provide evidence that answer-level evidence compression contributes to final answer selection.