arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16764cs.SE

RECTIFY:用于评估后RAG诊断、修复和验证的交互式工作台

RECTIFY: An Interactive Workbench for Post-Evaluation RAG Diagnosis, Repair, and Verification

  • University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Keerthana Murugaraj, Salima Lamsiyah, Martin Theobald

AI总结:

RECTIFY是一个交互式Streamlit工作台,将RAG评估结果转化为可审计的修复工作流,通过过滤、分类和生成可编辑修复卡片,帮助开发者做出可检查的修复决策。

AI中文摘要:

检索增强生成(RAG)评估器能够识别诸如检索薄弱、依据不足、回答不完整以及生成内容缺乏支持等失败,但它们很少帮助开发者决定下一步应修复什么。我们提出了RECTIFY,一个交互式Streamlit工作台,可将评估后的RAG案例转化为可审计的修复工作流。RECTIFY过滤掉无需修复的案例,将剩余的失败案例归类为可操作的问题系列和细粒度的修复切片,并生成可编辑的修复卡片,开发者可以通过沙箱重运行来批准、拒绝或验证这些卡片。在一个受控的RAG基准上,RECTIFY展示了跨BM25、稠密和混合检索的可解释失败概况:BM25主要触发噪声检索修复,而稠密和混合检索则留下较小规模的多部分检索不足和证据未充分利用的案例。进一步的分析表明,预过滤减少了不必要的修复候选,并且切片级路由比广泛的问题系列级诊断产生更具针对性的修复卡片。RECTIFY作为开源Streamlit工作台公开可用,旨在帮助开发者将评估结果转化为可检查的修复决策。

英文摘要:

Retrieval-Augmented Generation (RAG) evaluators can identify failures such as weak retrieval, poor grounding, incomplete answers, and unsupported generation, but they rarely help developers decide what to repair next. We present RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows. RECTIFY filters cases that do not require repair, routes remaining failures into actionable families and finegrained repair slices, and generates editable repair cards that developers can approve, reject, or verify through sandbox reruns. On a controlled RAG benchmark, RECTIFY surfaces interpretable failure profiles across BM25, dense, and hybrid retrieval: BM25 mainly triggers noisy-retrieval repairs, while dense and hybrid retrieval leave smaller sets of multi-part underretrieval and underused-evidence cases. Additional analyses show that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis. RECTIFY is publicly available as an open-source Streamlit workbench 1 for helping developers turn evaluation results into inspectable repair decisions.

↑