arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1216 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1216 篇

2601.09527 2026-01-15 cs.LG cs.AI cs.PF 57%

Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs

在消费级Blackwell GPU上进行私有LLM推理:为中小企业实现低成本本地部署的实用指南

Jonathan Knoop, Hendrik Holtmann

机构 * IE Business University(IE商学院)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文评估了消费级Blackwell GPU在LLM推理中的性能,证明其在多数中小企业工作负载中可替代云服务,但长上下文RAG任务仍需高端GPU。

Comments 15 pages, 18 tables, 7 figures. Includes link to GitHub repository and Docker image for reproducibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09523 2026-01-15 cs.IR 57%

TEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval

TEMPO:一个用于时间推理密集检索的多领域基准

Abdelrahman Abdallah, Mohammed Ali, Muhammad Abdul-Mageed, Adam Jatowt

专题命中 RAG评测 :RAG(abstract);分类 cs.IR

AI总结 TEMPO是一个结合时间推理与多领域检索的基准,通过复杂查询和多步骤检索规划,评估检索系统在时间信息处理上的挑战和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04416 2026-01-09 cs.AI 57%

Transitive Expert Error and Routing Problems in Complex AI Systems

复杂人工智能系统中的传递专家误差与路由问题

Forest Mars

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 复杂AI系统中,传递专家误差导致领域边界处因果错误输出,通过多专家激活、边界校准和覆盖检测等干预措施加以解决。

Comments 31pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02504 2026-01-07 cs.SE cs.AI cs.CY 57%

Enhancing Debugging Skills with AI-Powered Assistance: A Real-Time Tool for Debugging Support

利用AI赋能提升调试技能:一种实时的调试支持工具

Elizaveta Artser, Daniil Karol, Anna Potriasaeva, Aleksei Rostovskii, Katsiaryna Dzialets, Ekaterina Koshchenko, Xiaotian Su, April Yi Wang, Anastasiia Birillo

机构 * JetBrains Research(JetBrains研究)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出一种基于AI的实时调试辅助工具,通过减少LLM调用并提升准确性,帮助提升编程教育中的调试技能。

Comments Accepted at ICSE SEET 2026, 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12830 2025-12-25 cs.CY cs.AI 57%

Gobernanza y trazabilidad "a prueba de AI Act" para casos de uso legales: un marco técnico-jurídico, métricas forenses y evidencias auditables

AI Act合规的法律用例治理与可追溯性:技术-法律框架、司法计量学与可审计证据

Alex Dantart

机构 * Humanizing Internet

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出一个开源框架rag-forense,用于确保AI系统在法律领域符合欧盟AI法案的可验证合规性。

Comments in Spanish and English languages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19769 2025-12-24 cs.SE cs.AI cs.PL 57%

A Declarative Language for Building And Orchestrating LLM-Powered Agent Workflows

构建和协调基于大语言模型的智能体工作流的声明式语言

Ivan Daunis

机构 * PayPal

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出一种声明式语言,用于构建和协调基于大语言模型的智能体工作流,通过统一DSL减少开发时间和部署速度,提升非工程师的使用效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19253 2025-12-16 cs.IR 57%

DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research

DeepResearchGym: 一个免费、透明且可复现的深度研究评估沙盒

João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao, Abhijay Paladugu, Pranav Setlur, Jiahe Jin, Jamie Callan, João Magalhães, Bruno Martins, Chenyan Xiong

专题命中 RAG评测 :retriever(abstract);分类 cs.IR

AI总结 DeepResearchGym提供了一个免费、透明且可复现的深度研究评估沙盒,结合可复现的搜索API和严谨的评估协议,用于评估深度研究系统性能,展示其在成本效益训练中的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20209 2025-12-11 cs.LG cs.AI 57%

Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset

利用常规实验室数据评估早期癌症检测的可行性:对机器学习方法在不平衡数据集上的评估

Shumin Li

专题命中 RAG评测 :retriever(abstract);分类 cs.AI

AI总结 本研究评估了利用常规实验室数据通过机器学习检测癌症的可行性,发现其在临床分类上表现欠佳,需整合多模态数据以提升效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08998 2025-12-11 eess.IV cs.AI cs.CV 57%

DermETAS-SNA LLM: A Dermatology Focused Evolutionary Transformer Architecture Search with StackNet Augmented LLM Assistant

DermETAS-SNA LLM:一种专注于皮肤科的进化Transformer架构搜索与StackNet增强LLM助手

Nitya Phani Santosh Oruganty, Keerthi Vemula Murali, Chun-Kit Ngan, Paulo Bandeira Pinho

机构 * Data Science Program, Worcester Polytechnic Institute, Massachusetts, USA(沃斯特理工学院数据科学项目,马萨诸塞州,美国) Computer Science Department, Worcester Polytechnic Institute, Massachusetts, USA(沃斯特理工学院计算机科学系,马萨诸塞州,美国) PASE Advisory Group, New Jersey, USA(PASE顾问小组,新泽西州,美国)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 DermETAS-SNA LLM结合进化Transformer架构搜索与StackNet增强LLM,旨在提升皮肤病诊断的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07921 2025-12-10 cs.SE cs.AI 57%

DeepCode: Open Agentic Coding

DeepCode: 开放代理编程

Zongwei Li, Zhonghang Li, Zirui Guo, Xubin Ren, Chao Huang

机构 * The University of Hong Kong(香港大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 DeepCode通过系统信息流管理,实现文档到代码库的高保真合成,超越现有商业代理和人类专家,推动自主科学复现的发展。

Comments for source code, please see https://github.com/HKUDS/DeepCode

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04307 2025-12-05 cs.LG cs.AI 57%

Evaluating Long-Context Reasoning in LLM-Based WebAgents

评估基于大语言模型的WebAgent的长上下文推理能力

Andy Chung, Yichi Zhang, Kaixiang Lin, Aditya Rawal, Qiaozi Gao, Joyce Chai

机构 * University of Michigan(密歇根大学) Amazon(亚马逊)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文评估了基于大语言模型的WebAgent在长上下文场景中的推理能力,发现随着上下文长度增加,性能显著下降,提出隐式RAG方法以改进任务执行。

Comments Accepted NeurIPS 25 LAW Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01786 2025-12-02 cs.AI cs.LG 57%

Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems

谁评判法官?LLM陪审团即需:构建可信的LLM评估系统

Xiaochuan Li, Ke Wang, Girija Gouda, Shubham Choudhary, Yaqun Wang, Linwei Hu, Joel Vaughan, Freddy Lecue

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出LLM陪审团即需,通过动态学习框架提升LLM评估的可靠性和可扩展性,实验表明其在关键决策中的相关性优于传统方法。

Comments 66 pages, 22 figures, 37 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00991 2025-12-02 cs.CL 57%

Advancing Academic Chatbots: Evaluation of Non Traditional Outputs

推进学术聊天机器人:非传统输出的评估

Nicole Favero, Francesca Salute, Daniel Hardt

机构 * Copenhagen Business School(哥本哈根商学院)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL

AI总结 本研究评估了非传统学术输出生成,比较了两种检索策略并测试了LLMs在幻灯片和播客脚本生成中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00765 2025-11-17 cs.CL 57%

Decomposing and Revising What Language Models Generate

Zhichao Yan, Jiaoyan Chen, Jiapu Wang, Xiaoli Li, Ru Li, Jeff Z. Pan

机构 * School of Computer and Information Technology, Shanxi University, Taiyuan, China(计算机与信息科技学院,山西大学,太原,中国) Department of Computer Science, University of Manchester, Manchester, England(计算机科学系,曼彻斯特大学,曼彻斯特,英国) Beijing University of Technology, Beijing, China(北京理工大学,北京,中国) Singapore University of Technology and Design(新加坡科技与设计大学) ILCC, School of Informatics, University of Edinburgh, Edinburgh, England(ILCC,信息学院,爱丁堡大学,爱丁堡,英国)

专题命中 RAG评测 :retriever(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15045 2025-11-14 cs.CL 57%

Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework

Jinge Wu, Zhaolong Wu, Ruizhe Li, Tong Chen, Abul Hasan, Yunsoo Kim, Jason P. Y. Cheung, Teng Zhang, Honghan Wu

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06183 2025-11-11 cs.CL 57%

BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering

Ryuhei Miyazato, Ting-Ruen Wei, Xuyang Wu, Hsin-Tai Wu, Kei Harada

机构 * The University of Electro-Communications(电子通信大学) Santa Clara University(圣克拉拉大学) DOCOMO Innovations, Inc.(doCOMO创新公司)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26130 2025-11-06 cs.SE cs.AI cs.LG 57%

Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation

Musfiqur Rahman, SayedHassan Khatoonabadi, Emad Shihab

机构 * Concordia University(康科德大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

Comments Pre-print submitted for reviwer to TOSEM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14681 2025-11-04 cs.CL 57%

Large Language Models as Medical Codes Selectors: a benchmark using the International Classification of Primary Care

Vinicius Anjos de Almeida, Vinicius de Camargo, Raquel Gómez-Bravo, Egbert van der Haring, Kees van Boven, Marcelo Finger, Luis Fernandez Lopez

机构 * Medical School, University of São Paulo(圣保罗大学医学院) University of São Paulo(圣保罗大学) Department of Epidemiology, School of Public Health, University of São Paulo(圣保罗大学流行病学系) Rehaklinik, Centre Hospitalier Neuro-psychiatrique (CHNP)(康复诊所,神经精神病中心(CHNP)) Radboud University(拉德堡德大学) Department of Primary and Community Care, Radboud University(初级与社区护理系,拉德堡德大学) Institute of Mathematics and Statistics, University of Sao Paulo(数学与统计学研究所,圣保罗大学)

专题命中 RAG评测 :retriever(abstract);分类 cs.CL

Comments Accepted at NeurIPS 2025 as a poster presentation in The Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance (https://openreview.net/forum?id=Kl7KZwJFEG). 33 pages, 10 figures (including appendix), 15 tables (including appendix). To be submitted to peer-reviewed journal. For associated code repository, see https://github.com/almeidava93/llm-as-code-selectors-paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25820 2025-10-31 cs.AI cs.HC 57%

Symbolically Scaffolded Play: Designing Role-Sensitive Prompts for Generative NPC Dialogue

Vanessa Figueiredo, David Elumeze

机构 * ExplorAI, Department of Computer Science, University of Regina(ExplorAI,计算机科学系,里嘉大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14101 2025-10-24 cs.CL 57%

MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations

Ernests Lavrinovics, Russa Biswas, Katja Hose, Johannes Bjerva

机构 * Department of Computer Science, Aalborg University(计算机科学系,奥胡斯大学) Institute of Logic and Computation, TU Wien(逻辑与计算研究所,维也纳技术大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20934 2025-10-17 cs.SE cs.AI 57%

Leveraging LLMs, IDEs, and Semantic Embeddings for Automated Move Method Refactoring

Abhiram Bellur, Fraol Batole, Mohammed Raihan Ullah, Malinda Dilhara, Yaroslav Zharov, Timofey Bryksin, Kai Ishikawa, Haifeng Chen, Masaharu Morimoto, Shota Motoura, Takeo Hosomi, Tien N. Nguyen, Hridesh Rajan, Nikolaos Tsantalis, Danny Dig

机构 * University of Colorado(科罗拉多大学) Tulane University(路易斯安那州立大学) Amazon Web Services(亚马逊网络服务) JetBrains Research(JetBrains研究) NEC Corporation(日本电报电话公司) NEC Laboratories America(日本电报电话美洲实验室) University of Texas at Dallas(德克萨斯大学达拉斯分校) Concordia University(康科迪亚大学) University of Colorado, JetBrains Research(科罗拉多大学,JetBrains研究)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

Comments Published at the International Conference on Software Maintenance and Evolution (ICSME'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08648 2025-10-13 cs.LG cs.AI 57%

Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity

Edward Y. Chang, Ethan Y. Chang

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

Comments 24 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07920 2025-10-10 cs.AI 57%

Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents

Xiangyu Li, Yawen Zeng, Xiaofen Xing, Jin Xu, Xiangmin Xu

机构 * South China University of Technology(华南理工大学) Pazhou Lab(黄埔实验室)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18178 2025-10-02 cs.AI cs.CE cs.LG 57%

Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM

Ling Yue, Nithin Somasekharan, Tingwen Zhang, Yadi Cao, Shaowu Pan

机构 * Rensselaer Polytechnic Institute(拉特格斯理工学院) University of California San Diego(加州大学圣地亚哥分校)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22678 2025-10-02 cs.CL 57%

Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions

Mohammad Almansoori, Komal Kumar, Hisham Cholakkal

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL

Comments 14 page, 4 figures, 61 references, presented in MICCAI (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25197 2025-10-01 cs.SE cs.AI cs.PL 57%

Towards Repository-Level Program Verification with Large Language Models

Si Cheng Zhong, Xujie Si

机构 * University of Toronto(多伦多大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

Comments Accepted to LMPL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21969 2025-09-29 cs.RO cs.AI 57%

DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation

Tianjun Gu, Linfeng Li, Xuhong Wang, Chenghua Gong, Jingyu Gong, Zhizhong Zhang, Yuan Xie, Lizhuang Ma, Xin Tan

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19218 2025-09-24 cs.CV cs.AI 57%

HyKid: An Open MRI Dataset with Expert-Annotated Multi-Structure and Choroid Plexus in Pediatric Hydrocephalus

Yunzhi Xu, Yushuang Ding, Hu Sun, Hongxi Zhang, Li Zhao

机构 * College of Biomedical Engineering and Instrument Science, Zhejiang University, Hangzhou, China(浙江大学生物医学工程与仪器科学学院) Department of Radiology, Children’s Hospital, Zhejiang University School of Medicine, National Clinical Research Center for Child Health, Hangzhou, China(浙江大学医学院儿童医院放射科) Neurosurgery of Zhejiang Hospital, Hangzhou, China(浙江医院神经外科)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.AI

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18868 2025-09-24 cs.AI 57%

Memory in Large Language Models: Mechanisms, Evaluation and Evolution

Dianxing Zhang, Wendong Li, Kani Song, Jiaye Lu, Gang Li, Liuchun Yang, Sheng Li

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

Comments 50 pages, 1 figure, 8 tables This is a survey/framework paper on LLM memory mechanisms and evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06412 2025-09-09 cs.DL cs.IR 57%

Compare: A Framework for Scientific Comparisons

Moritz Staudinger, Wojciech Kusa, Matteo Cancellieri, David Pride, Petr Knoth, Allan Hanbury

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.IR

Comments Accepted at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏