DS@GT ARC参加LongEval:科学问答中的引用完整性和事实基础
DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究在科学问答中传统评估指标与引用完整性的差异,通过Corrective RAG和CiteFix构建纠正管道,对比前沿模型,发现前沿模型答案生成不依赖文档上下文,而纠正管道提升了引用忠实度和答案基础,提出需奖励严格答案基础的评估指标。
AI中文摘要:
本文描述了DS@GT ARC提交给2026年CLEF LongEval任务4关于检索增强生成(RAG)的内容。我们研究了传统自然语言评估指标与应用于RAG问答系统的引用完整性之间的差异。使用Corrective RAG(CRAG)和CiteFix评估一个纠正管道,对比基线和前沿模型基准RAG问答分数。前沿模型最大化了答案相关性和流畅性分数,但我们的诊断表明前沿模型在生成答案时未使用文档上下文就能正确识别相关文档。而我们的纠正管道通过预生成过滤块和生成后对引用材料严格强制蕴含,略微提高了引用忠实度和答案基础。我们提出,对可信RAG问答的评估需要奖励严格答案基础的指标。
英文摘要:
This paper describes DS@GT ARC's submission to the CLEF 2026 LongEval Task 4 on Retrieval-Augmented Generation (RAG). In this submission, we examine a divergence between traditional natural language evaluation metrics and citation integrity as applied to RAG QA systems. We evaluate a corrective pipeline using Corrective RAG (CRAG) and CiteFix against baseline and frontier model benchmark RAG QA scores. While frontier models maximized answer relevance and fluency scores, our RAGAs LLM-as-judge diagnostics indicate that frontier models would correctly identify relevant documents without using their context in answer generation. Conversely, by filtering chunks pre-generation and enforcing strict entailment of generated claims to the cited material post-generation, our corrective pipeline marginally improved citation faithfulness and answer grounding. We propose that evaluation of trustworthy RAG QA requires metrics that reward strict answer grounding.