arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-02-03 至 2026-02-03 共收录 33 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 6 篇

2509.16584 2026-02-03 cs.CL cs.AI 62%

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00723 2026-02-03 cs.LG cs.AI 57%

Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity

重新思考幻觉:正确性、一致性与提示多样性

Prakhar Ganesh, Reza Shokri, Golnoosh Farnadi

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出提示多样性框架,揭示LLM幻觉评估中一致性的重要性,指出现有检测与缓解技术的局限性。

Comments To appear at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00641 2026-02-03 cs.AI 57%

AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents

AgentAuditor: 为LLM代理提供人类水平的安全性和安全性评估

Hanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li, Guibin Zhang, Kun Wang, Tongliang Liu, Hanan Salam

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 AgentAuditor通过构建经验记忆和多阶段检索生成过程,提升LLM在代理安全性和安全性评估中的性能,达到人类水平的准确性。

Comments This paper is accepted by 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏