arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-02-03 至 2026-02-03 共收录 6 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 6 篇

2602.00010 2026-02-03 cs.IR cs.AI cs.CL 75%

ChunkNorris: A High-Performance and Low-Energy Approach to PDF Parsing and Chunking

ChunkNorris: 一种高性能且低能耗的PDF解析与分块方法

Mathieu Ciancone, Clovis Varangot-Reille, Marion Schaeffer

机构 * Wikit Laboratoire Hubert Curien(Hubert Curien实验室) Université Jean Monnet(Jean Monnet大学) INSA Rouen Normandie(Rouen Normandie国立理工学院)

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 ChunkNorris通过高效启发式方法实现PDF解析与分块的高性能低能耗,优于现有技术,适用于资源受限的检索增强生成任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15567 2026-02-03 cs.CR cs.SE 67%

MalCVE: Malware Detection and CVE Association Using Large Language Models

MalCVE: 使用大语言模型进行恶意软件检测与CVE关联

Eduard Andrei Cristea, Petter Molnes, Jingyue Li

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

AI总结 MalCVE利用大语言模型检测恶意软件并关联CVE,实现高准确率的恶意软件识别和漏洞关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01885 2026-02-03 cs.CL cs.AI 62%

ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support

ES-MemEval:在个性化长期情感支持中评估对话代理的基准测试

Tiantian Chen, Jiaqi Lu, Ying Shen, Lin Zhang

机构 * Tongji University(同济大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

AI总结 ES-MemEval通过评估长期情感支持中的记忆能力,揭示了显式长期记忆对减少幻觉和提升个性化的重要性,同时指出了检索增强模型在时间动态方面的局限性。

Comments 12 pages, 7 figures. Accepted to The Web Conference (WWW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16584 2026-02-03 cs.CL cs.AI 62%

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00723 2026-02-03 cs.LG cs.AI 57%

Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity

重新思考幻觉:正确性、一致性与提示多样性

Prakhar Ganesh, Reza Shokri, Golnoosh Farnadi

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出提示多样性框架,揭示LLM幻觉评估中一致性的重要性,指出现有检测与缓解技术的局限性。

Comments To appear at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00641 2026-02-03 cs.AI 57%

AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents

AgentAuditor: 为LLM代理提供人类水平的安全性和安全性评估

Hanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li, Guibin Zhang, Kun Wang, Tongliang Liu, Hanan Salam

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 AgentAuditor通过构建经验记忆和多阶段检索生成过程,提升LLM在代理安全性和安全性评估中的性能,达到人类水平的准确性。

Comments This paper is accepted by 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏