arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12190cs.SEcs.AI

面向科学代码理解的检索增强生成

Retrieval-Augmented Generation for Scientific Code Understanding

  • ETH Zürich(苏黎世联邦理工学院)
  • Paul Scherrer Institute(保罗·谢尔研究所)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Aaron Nobile, Andreas Adelmann, Mohsen Sadr

AI总结:

本研究构建基于小型开源模型的本地RAG系统,通过离线摄取与在线回答分离,在科学代码理解基准上9B模型取得最高分,证明前端加载代码理解可实现隐私保护的本地编码智能体。

AI中文摘要:

大型语言模型已成为现代编码助手的核心,但诸如Claude Code或Codex等最先进的系统依赖于非常庞大的云端托管模型,这带来了显著的计算成本和数据隐私问题。本研究探讨是否可以通过将计算负担从推理阶段转移,围绕小型开源模型构建一个实用且完全本地的编码智能体。我们开发了一个用于科学代码理解的检索增强生成(RAG)系统,该系统严格区分了昂贵的离线摄取阶段(包括解析、结构图构建、LLM生成的实体解释和嵌入)与轻量级的在线回答阶段。该系统在一个包含100个问题、涵盖11个类别的基准测试上进行了评估,该基准测试基于用C++编写的IPPL科学代码库,答案由独立的领先模型作为评判者进行评分。在七个回答模型中,我们发现模型家族和检索质量比参数数量更重要,即一个9B模型取得了最高平均分(0.795),优于我们流程中的更大模型以及嵌入在Claude Code检索架构中的相同模型。结果表明,将代码理解前置到可重用的、针对代码库特化的向量存储中,能使小型本地模型提供基于事实且针对仓库的答案,使该智能体非常适合作为内部科学代码库的隐私保护开发工具。

英文摘要:

Large language models have become central to modern coding assistants, but state-of-the-art systems such as Claude Code or Codex rely on very large, cloud-hosted models with significant computational cost and data-privacy implications. This work investigates whether a useful, fully local coding agent can be built around small open-source models by shifting the computational burden away from inference. We develop a Retrieval-Augmented Generation (RAG) system for scientific code understanding that strictly separates an expensive offline ingestion stage parsing, structural graph construction, LLM-generated entity explanations, and embedding from a lightweight online answering stage. The system is evaluated on a 100-question benchmark spanning eleven categories over the IPPL scientific codebase written in C++, with answers scored by an independent frontier model as the judge. Across seven answering models, we find that model family and retrieval quality matter more than parameter count, i.e. a 9B model achieves the highest average score (0.795), outperforming both larger models within our pipeline and the same models embedded in the Claude Code retrieval architecture. The results indicate that front-loading code understanding into a reusable, codebase-specialised vector store enables small local models to deliver grounded and repository-specific answers, making the agent well suited as a privacy-preserving development tool for in-house scientific codebases.

↑