弥合静态与智能体RAG在台湾历史问答中的应用
Bridging Static and Agentic RAG for Taiwanese Historical Question Answering
- National Taiwan University(国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过受控比较发现,智能体与静态RAG在台湾历史问答中总体性能相近但问题级差异显著,并引入事后选择器利用互补性,恢复神谕提升空间的60.34%。
AI中文摘要:
智能体检索增强生成(RAG)使语言模型能够根据先前检索到的证据调整检索策略,但尚不清楚这种自适应编排是否始终优于设计良好的静态流水线。我们针对台湾历史问答进行了智能体RAG与静态RAG的受控比较,两者共享相同的生成器和混合检索后端。尽管总体性能相似,但两种流水线在70.83%的问题上存在差异,其优势在平均后基本相互抵消。一个为每个问题选择更优响应的神谕(oracle)将综合得分比更优的单一流水线提高了0.2417,揭示了问题级选择存在巨大提升空间。因此,我们引入了一个事后选择器,用于比较两个响应及其引用的证据,其性能显著优于任一单一流水线,并恢复了神谕提升空间的60.34%。这些结果表明,总体比较可能掩盖检索策略之间有意义的问级差异,提示利用它们的互补性可能比寻求普遍优越的流水线更有成效。
英文摘要:
Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed static pipelines. We conduct a controlled comparison of agentic and static RAG for Taiwanese historical question answering, sharing the same generator and hybrid retrieval backend. Despite similar aggregate performance, the two pipelines differ on 70.83% of questions, with their advantages largely canceling out when averaged. An oracle that selects the better response per question improves the composite score by 0.2417 over the better individual pipeline, revealing substantial headroom for question-level selection. We therefore introduce a post-hoc selector that compares the two responses and their cited evidence, significantly outperforming either individual pipeline and recovering 60.34% of the oracle headroom. These results show that aggregate comparisons can obscure meaningful question-level differences between retrieval strategies, suggesting that exploiting their complementarity may be more fruitful than seeking a universally superior pipeline.