arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你是在合成还是回忆?评估LLM在算法代码检索上的表现

Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval

Nickil Maveli, Antonio Vergari, Shay B. Cohen

arXiv 2610.02438首次发表:更新:

发表机构

University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AlgoREval基准,评估LLM在经典算法代码检索中的能力,发现准确率因语言和输入表示而异,提示增强和微调可提升性能,并警示需验证AI生成代码。

AI 中文摘要

大型语言模型(LLMs)在代码生成方面展现了强大的性能,其成功既依赖于回忆相关的算法知识,也依赖于推理如何应用这些知识。然而,现有的LLM流程是不透明的,没有明确区分这两个组成部分。我们认为,对于预训练语料库中广泛可获取的经典算法,其代码生成更适合被衡量为参数化代码检索:从内部化知识中重现一个命名算法,而非合成一个新算法。我们引入了AlgoREval,一个包含599个问题的基准测试,涵盖14个领域的77个经典算法、7种编程语言和4种图输入表示,以独立评估这一能力,并在零样本设置下评估了15个模型(7B-34B参数)。我们发现,即使对于广泛记录的算法,检索准确率在不同语言和输入表示之间也存在显著差异,并表明通过检索到的代码片段或结构化算法提示进行提示增强可以提高复杂算法的准确率,而SFT在语言上实现了更广泛的提升,GRPO在特定语言上实现了更大的每语言提升。总之,我们的结果将参数化代码检索确立为一种独特、可测量的能力,并警示在缺乏系统性验证的情况下部署AI生成的算法代码。代码和数据集可在以下网址获取:此https URL

英文摘要

Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B--34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnote{Code and dataset are available at https://github.com/Nickil21/AlgoREval

Comments30 pages (preprint)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑