面向基于证据的多跳问答中选择性回答的证据充分性边界学习
Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
AI总结:
该研究针对多跳问答的选择性回答问题,提出证据充分性边界训练框架,在Qwen2.5-3B-Instruct适配后,提升了边界定位与无依据答案率表现,同时保持问答效用。
AI中文摘要:
基于证据的问答系统应仅在提供的证据支持答案时才给出回答。在多跳问答中,这一要求难以实现,因为部分证据会让无依据的答案显得合理。我们通过证据充分性边界研究选择性回答:对于同一问题,模型应在无依据或部分支持的上下文下弃权(不执行),在上下文首次变得充分时给出答案,且在添加冗余证据时保持答案稳定。我们提出证据充分性边界训练,这是一种原生生成式训练框架,可构建有序证据链并直接监督从弃权到回答的转变。该方法结合了层级监督、边界翻转裕度、边界后稳定性及答案召回保护。我们从HotpotQA、2WikiMultiHopQA和MuSiQue构建证据链,随后在外部不可回答数据集上,采用链指标、原始问答效用及无依据答案率评估模型。使用Qwen2.5-3B-Instruct和LoRA适配时,证据充分性边界训练在测试系统中实现最强的边界定位,翻转准确率达0.807,而令牌级弃权基线为0.781;其在外部不可回答评估中还实现最低的整体无依据答案率,为0.095,而同基线为0.101,同时保持有竞争力的原始问答F1。结果表明,当训练标记出拒绝应转向回答的证据层级时,基于证据的选择性回答性能会提升。
英文摘要:
Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible. We study selective answering through evidence sufficiency boundaries: for the same question, a model should abstain under unsupported or partially supported context, answer when the context first becomes sufficient, and keep the answer stable when redundant evidence is added. We introduce Evidence Sufficiency Boundary Training, a generation-native training framework that constructs ordered evidence chains and supervises the abstain-to-answer transition directly. The method combines level supervision, a boundary flip margin, post-boundary stability, and answer recall protection. We build evidence chains from HotpotQA, 2WikiMultiHopQA, and MuSiQue, then evaluate models with chain metrics, raw QA utility, and unsupported-answer rates on external non-answerable sets. With Qwen2.5-3B-Instruct and LoRA adaptation, Evidence Sufficiency Boundary Training gives the strongest boundary localization among the tested systems, with flip accuracy of 0.807 compared with 0.781 for a token-level abstention baseline. It also achieves the lowest overall unsupported-answer rate on external non-answerable evaluation, 0.095 compared with 0.101 for the same baseline, while retaining competitive raw QA F1. The results show that grounded selective answering improves when training marks the evidence level where refusal should give way to answering.