arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-09-01 至 2026-09-01 共收录 11 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 图谱与结构化RAG 11 篇

2608.29290 2026-09-01 cs.IR cs.SE 新提交 91%

Database-Augmented RAG for Automated Repair of REST API Misuses

数据库增强的检索增强生成(RAG)用于REST API误用的自动修复

Shoei Inoue, Norihiro Yoshida, Erina Makihara, Shiyu Yang, Katsuro Inoue

专题命中 图谱与结构化RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 本研究构建11种不同数据库结构的RAG配置,对比基线方法,发现按版本和内容类型组织API规范的四数据库RAG方法,能将REST API误用修复率从54.3%提升至88.6%。

Comments Accepted at the 26th International Conference on Software Quality, Reliability, and Security (QRS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24921 2026-09-01 cs.AI 版本更新 90%

post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis

post-graph-rag:一款原生PostgreSQL图检索增强生成引擎

Chandan Rajah

专题命中 图谱与结构化RAG :RAG(title,title_cn);分类 cs.AI

AI总结 post-graph-rag是一款开源PostgreSQL原生图RAG引擎,解决现有图RAG的基础设施、图质量、时间累积问题,在三语料库对比中构建更优图且支持时间演化特性

Comments 31 pages, 6 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28611 2026-09-01 cs.CL cs.AI cs.CY 新提交 85%

Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

Gurukul AI:面向印度教育体系的交互式AI驱动教育平台

Isha Narang, Sneh Gosai, Mayank Singh

机构 * Indian Institute of Technology Gandhinagar(印度甘地讷格尔印度理工学院) Pandit Deendayal Energy University(潘迪特·迪恩代亚尔能源大学)

专题命中 图谱与结构化RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究针对现有AI教育工具不适配印度教育体系的问题,构建了NCERT大纲对齐的问答数据集,微调LLaMA 3.1 8B模型并部署RAG框架,推出交互式教育平台GurukulAI,缩小了全球LLM与印度教育需求的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04226 2026-09-01 cs.CL cs.AI cs.CY cs.IR cs.LG 83%

What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models

知识多样性与大语言模型中的知识崩溃

Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christensen, Chan Young Park, Isabelle Augenstein

机构 * University of Copenhagen(哥本哈根大学) Aalborg University Copenhagen(奥胡斯大学哥本哈根分校) Stanford University(斯坦福大学) University of Colorado Boulder(科罗拉多大学波德分校) Microsoft Research(微软研究院)

专题命中 图谱与结构化RAG :RAG(summary_cn,abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本研究通过评估大语言模型的知识多样性,发现模型大小影响多样性,RAG技术提升多样性,但文化背景影响显著,且国家特定信息更偏向英语。

Comments Accepted to EMNLP 2026; 18 pages, 8 figures, 4 tables; v2: Fixed the modeling for table 3, random effect is the model version; v3: Fixed minor formatting issues in tables 2 and 3; v4: Fixed some typos and model description; v5: Updated metadata; v6: Improved search baseline, writing revisions, added comparisons to semantic similarity only approaches; v7 changelog: EMNLP camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24888 2026-09-01 cs.AI 版本更新 77%

SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Planning

SIMGUIDE:用于个性化智能体规划的过程式基础多上下文表示

Chirag Shah

机构 * University of Washington(华盛顿大学)

专题命中 图谱与结构化RAG :RAG(summary_cn,abstract_cn);分类 cs.AI

AI总结 本研究针对个性化AI智能体无法适配用户多场景优先级的问题,提出SIMGUIDE方法,构建SIMBENCH基准验证,发现过程式基础的Sims优于RAG,且表示格式是关键设计变量。

Comments Previous version had some results that were hard to verify or reproduce, so new experiments were done and many parts of the paper were changed to reflect that

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.31082 2026-09-01 cs.AI cs.CL cs.DB 新提交 75%

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

通过非结构化数据的自适应结构化实现 token 高效的数据推理智能体

Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos

机构 * Harvard University(哈佛大学)

专题命中 图谱与结构化RAG :RAG(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.DB

AI总结 该研究提出智能体数据拆解方法,通过自适应推测性结构化非结构化数据,降低 LLM 智能体推理成本,在 FanOutQA 基准上实现 53% 的成本削减且保持准确率,为智能体推理数据基础设施提供了初步方案。

Comments 7 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30947 2026-09-01 cs.CL 版本更新 70%

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

将人工智能研究扩展到人文学科:一个用于证据基础学术的多智能体框架

Yating Pan, Jiajun Zhang, Jun Wang, Qi Su

机构 * Department of Information Management(信息管理系) Research Center for Digital Humanities(数字人文研究中心) School of Foreign Languages(外国语言学院) Institute for Artificial Intelligence(人工智能研究院)

专题命中 图谱与结构化RAG :RAG(abstract,abstract_cn);分类 cs.CL

AI总结 提出SPIRE多智能体框架,通过将人文学科操作建模为协作智能体角色,结合多尺度细读检索,实现基于证据的论证,在古典文献基准上优于现有方法。

Comments Accepted to the Main Conference of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28608 2026-09-01 cs.CL cs.AI cs.IR 新提交 67%

NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

NLP驱动的古印度翻译医学文本知识提取与主题分类

M. S. Rajeevan, B. Mini Devi, V. S. Anoop, C. Mallikarjuna

专题命中 图谱与结构化RAG :knowledge retrieval(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本研究采用NLP相关技术对古印度翻译医学文本进行知识提取与主题分类,实现阿育吠陀医学智慧的计算组织,提升其可获取性,推动多学科发展。

Comments 19 pages, 8 figures, 5 tables. Presented at the National Conference on "Reimagining LIS Education: Integrating Indian Knowledge Systems with NEP 2020" (March 2025), organized by Tata Institute of Social Sciences (TISS) and the Indian Association of Teachers of Library and Information Science (IATLIS). Recipient of the Best Paper Award

Journal ref In Proc. TISS-IATLIS National Conference 2025: Reimagining LIS Education: Integrating Indian Knowledge Systems with NEP 2020, Vol. 1, p. 351, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13111 2026-09-01 cs.CL cs.AI cs.IR 版本更新 67%

CORE-T: COherent REtrieval of Tables for Text-to-SQL

CORE-T: 面向文本到SQL的表格连贯检索

Hassan Soliman, Vivek Gupta, Dan Roth, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(普适知识处理实验室(UKP实验室),计算机科学系 TU Darmstadt 和应用网络安全国家研究中心 ATHENE,德国) Arizona State University(亚利桑那州立大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 图谱与结构化RAG :dense retrieval(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 提出CORE-T框架,通过LLM生成元数据和预计算兼容性缓存,在无需训练的情况下从异构表集合中高效检索连贯可连接的表集合,提升表选择F1最多22.7点并减少40%的表数量。

Comments Accepted at EMNLP Main 2026. Camera-ready version with revisions following peer review. Code and data available at: https://github.com/UKPLab/emnlp2026-core-t

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29345 2026-09-01 cs.AI cs.CL 新提交 62%

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History:带有细粒度知识标注的历史驱动型Text-to-SQL基准测试

Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu

机构 * Zhejiang University(浙江大学)

专题命中 图谱与结构化RAG :retriever(abstract);分类 cs.CL、cs.AI

AI总结 本研究推出带细粒度知识标注的BIRD-History基准测试,针对现有Text-to-SQL系统无法利用历史查询日志处理隐含领域知识的问题,提出插件式检索器,可提升多个Text-to-SQL系统的未明确查询处理性能。

Comments Accepted at Findings of the Association for Computational Linguistics: EMNLP, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28978 2026-09-01 cs.AI 新提交 57%

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

选择性遗忘:面向长期大语言模型智能体的基于图的记忆框架

Theo Rusu, Sourena Khanzadeh, Manar Alalfi

机构 * Toronto Metropolitan University(多伦多都会大学) The Creative School(创意学院) Flybits(弗莱比特公司)

专题命中 图谱与结构化RAG :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本研究评估了基于图的长期LLM智能体记忆框架的假设,发现其在LongMemEval上未优于扁平向量基线,但遗忘模块能高效剪枝节点且性能损失小,结果受提取器和单一基准限制。

详情

展开后加载摘要…

URL PDF HTML 收藏