AI 中文总结
该研究推出LayerRAG-Bench跨层可靠性基准,含多领域任务与模型数据,发现模式标准化无法修复部分故障,验证了分层评估原则的重要性。
AI 中文摘要
智能体检索增强生成系统生成的答案看似有依据,却可能在证据层、工具契约层、授权层或会话状态层失效。我们推出LayerRAG-Bench,这是一个受控的跨层可靠性基准,涵盖8个企业领域、240个任务、9种故障场景、2种契约模式,以及来自OpenAI、Anthropic和Gemini的9个模型的38880条实时任务级记录。模式标准化将模式漂移成功率从0.000提升至0.913,但模式标准化无法修复过时证据、缺失工具输出、权限被拒和错误会话上下文问题。仅基于依据性的评估在过时证据和错误会话证据下也会产生大量误报。这些结果支持分层评估原则:可靠性干预措施应仅因修复其目标层而被认可,而非被视为通用解决方案。
英文摘要
Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are not recovered by schema normalization. Groundedness-only evaluation also produces substantial false positives under stale and wrong-session evidence. These results support a layer-specific evaluation principle: a reliability intervention should be credited for repairing its target layer without being mistaken for a universal fix.
Comments10 pages, 9 tables. Code and data: https://github.com/MusaShams/layerrag-bench