arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

2026-01-08 至 2026-01-08 共收录 8 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码生成 3 篇

2601.03878 2026-01-08 cs.SE 79%

Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design

通过LLM理解基于规范的代码生成:一项实证研究设计

Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles, Jesús M. González-Barahona

专题命中 代码生成 :code generation(title,abstract);分类 cs.SE

AI总结 本文通过实证研究设计,探讨人类干预在基于规范的LLM代码生成过程中对代码质量和动态的影响。

Comments This paper is a Stage 1 Registered Report. The study protocol and analysis plan were peer reviewed and accepted at SANER 2026 with a Continuity Acceptance (CA) score for Stage 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03780 2026-01-08 cs.SE 79%

Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study

基于编程语言知识单元(KUs)评估和改进代码生成基准的代表性——一项实证研究

Md Ahasanuzzaman, Bram Adams, Emad Fallahzadeh, Gustavo A. Oliva, Ahmed E. Hassan

专题命中 代码生成 :code generation(title,abstract);分类 cs.SE

AI总结 本文通过分析编程语言知识单元(KUs)的覆盖情况,提出基于提示的框架改进代码生成基准的代表性,从而更准确评估LLM的代码生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03640 2026-01-08 cs.SE cs.CR 79%

Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test

LLM代码生成中的verbatim数据转录失败:一种状态跟踪压力测试

Mohd Ariful Haque, Kishor Datta Gupta, Mohammad Ashiqur Rahman, Roy George

专题命中 代码生成 :code generation(title,abstract);分类 cs.SE

AI总结 本文提出了一种最小化的基准测试,用于评估LLM在生成代码时对verbatim数据转录的可靠性,通过精确字符串包含和状态跟踪分析来检测长周期生成失败。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 软件智能体 1 篇

2601.04171 2026-01-08 cs.LG 57%

Agentic Rubrics as Contextual Verifiers for SWE Agents

代理 rubrics 作为 SWE 代理的上下文验证器

Mohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong He

专题命中 软件智能体 :repository(abstract);分类 cs.LG

AI总结 代理 rubrics 通过上下文感知的检查清单提升 SWE 代理的验证效率与准确性。

Comments 31 pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 代码评测 1 篇

2601.03432 2026-01-08 cs.SE 57%

CodeEval: A pedagogical approach for targeted evaluation of code-trained Large Language Models

CodeEval: 一种面向代码训练大语言模型的教育性评估方法

Danny Brahman, Mohammad Mahoor

专题命中 代码评测 :code generation(abstract);分类 cs.SE

AI总结 CodeEval通过多维基准数据集和开源执行框架,针对代码训练大语言模型的评估与改进提供教育性方法。

Comments Accepted at the International Joint Conference on Natural Language Processing & Asia-Pacific Chapter of the Association for Computational Linguistics, 2025. Will be published at ACL anthology

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 仓库级理解 3 篇

2507.05281 2026-01-08 cs.SE cs.CL 81%

CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks

CoreCodeBench:通过细粒度仓库级任务解耦代码智能

Lingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan, Yaoming Zhu, Jun Xu, Zongyu Wang, Lin Qiu, Xunliang Cai, Xuezhi Cao, Weiwen Liu, Weinan Zhang, Yong Yu

机构 * Shanghai Jiao Tong University(上海交通大学) Meituan(美团)

专题命中 仓库级理解 :repository(title,abstract);分类 cs.SE、cs.CL

AI总结 CoreCodeBench通过细粒度仓库级任务解耦编码能力,揭示LLM在编程任务中的非单一性,提供更精确的评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05405 2026-01-08 astro-ph.HE astro-ph.IM 78%

The Open mulTiwavelength Transient Event Repository (OTTER): Infrastructure Release and Tidal Disruption Event Catalog

开放多波段暂现事件库(OTTER):基础设施发布与潮汐破坏事件目录

Noah Franz, Kate D Alexander, Sebastian Gomez, Collin T Christy, Tanmoy Laskar, Sjoert van Velzen, Nicholas Earl, Suvi Gezari, Mitchell Karmen, Raffaella Margutti, Jeniveve Pearson, V. Ashley Villar, Ann I Zabludoff

专题命中 仓库级理解 :repository(title,abstract)

AI总结 OTTER是一个开放的多波段暂现事件数据库,旨在通过存储和分析多波段光度学数据,提高对暂现事件物理机制的理解,特别关注潮汐破坏事件的光度学档案。

Comments Accepted to ApJ. The OTTER web interface is available at https://otter.idies.jhu.edu and the API documentation (including example python notebooks demonstrating usage) is available at https://astro-otter.readthedocs.io. Please submit any feedback as issues on GitHub at https://github.com/astro-otter/otter/issues/new/choose

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16242 2026-01-08 cs.SE cs.DL 57%

Code Contribution and Credit in Science

科学中的代码贡献与信用

Eva Maxfield Brown, Isaac Slaughter, Nicholas Weber

专题命中 仓库级理解 :repository(abstract);分类 cs.SE

AI总结 研究探讨了科学领域中代码贡献与传统作者身份信用之间的关系,发现代码贡献者常未获认可,且高频编码与学术影响力指标呈负相关。

Comments Revisions after peer-review. This is the "Accepted" version of the paper!

详情

展开后加载摘要…

URL PDF HTML 收藏