Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
专题命中 代码评测 :code generation(abstract);program repair(abstract);分类 cs.SE
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
专题命中 代码评测 :code generation(abstract);program repair(abstract);分类 cs.SE
专题命中 代码评测 :code generation(abstract);code model(abstract);分类 cs.SE
机构 * MacPaw
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.LG
Comments Accepted to FORGE'25 Benchmarking on 15.01.2025, to be published by IEEE under the CC BY-NC-ND 4.0 license. This is the accepted version of the article (5 pages, 2 figures, 1 table). DOI will be added upon publication
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Accepted as Spotlight at the 42nd International Conference on Machine Learning (ICML 2025)
机构 * ByteDance(字节跳动)
专题命中 代码评测 :code generation(abstract);coding agent(abstract);分类 cs.AI
Comments 28 pages, 15 figures
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.CL、cs.AI
Comments ICLR 2025 Camera Ready
专题命中 代码评测 :code generation(abstract);program repair(abstract);分类 cs.SE
Comments This is a preprint submitted to ACM Transactions on Software Engineering and Methodology (TOSEM)
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Working in progress
专题命中 代码评测 :code generation(abstract);repository(abstract);分类 cs.LG
专题命中 代码评测 :program repair(abstract);unit test generation(abstract);分类 cs.SE
Comments 5 pages, 4 figures
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments 23 pages, 9 figures
专题命中 代码评测 :program synthesis(abstract);分类 cs.SE、cs.CL、cs.AI
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.AI、cs.LG
Comments There are some flaws in our experiments, we would like to fix it and publish a fixed version again in the very near future
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.AI、cs.LG
Comments ACL 2022 Camera-Ready
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, published at ICLR 2021
专题命中 代码评测 :program synthesis(abstract);分类 cs.SE、cs.AI、cs.LG
专题命中 代码评测 :repository(abstract,comments);分类 cs.CL、cs.AI、cs.LG
Comments 39 pages, repository at https://github.com/GEM-benchmark/NL-Augmenter
CORE-Bench:智能体编程时代代码检索的综合基准
专题命中 代码评测 :repository(abstract);coding agent(abstract)
AI总结 针对智能体编程中需求驱动的仓库级代码检索问题,构建了包含18万查询和10.6万上下文相关性标签的三级基准CORE-Bench,实验表明现有嵌入模型性能显著下降,微调可提升效果。
Comments Accepted By EMNLP 2026 Main
SkillNet: 创建、评估和连接AI技能
机构 * Zhejiang University(浙江大学) ; Tongji University(同济大学) ; Southeast University(东南大学) ; Alibaba Group(阿里巴巴集团) ; Tencent(腾讯) ; Fudan University(复旦大学) ; The University of Edinburgh(爱丁堡大学) ; Monash University(墨尔本大学) ; National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; Huzhou University(湖州大学) ; Hornor Device Co., Ltd(Hornor设备有限公司) ; Hangzhou Institute for Advanced Study, UCAS(杭州先进研究所,UCAS)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 SkillNet通过统一的本体和多维评估机制,大规模创建、评估和连接AI技能,提升代理性能并促进技能的持久掌握。
Comments http://skillnet.openkg.cn/; add SkillNet-Gym, a benchmark for evaluating skill retrieval, utilization, composition, and SkillNet-Fabric for task-specific skill routing through lightweight Wikis
Text-ADBench:基于大语言模型嵌入的文本异常检测基准
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本研究构建了基于LLM嵌入的文本异常检测基准,通过多模型、多数据集的实验发现嵌入质量决定检测效果,深度学习方法在LLM嵌入下无性能优势,还提出了高效评估策略并开源工具包。
AI4AI-Bench:用于递归自我改进算法设计中大语言模型智能体的基准测试
机构 * Navers Lab(Navers实验室) ; Tsinghua University(清华大学)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出AI4AI-Bench基准测试,测试大语言模型智能体设计训练算法的能力,实验显示现有智能体仅缩小了已有算法与最优值间不到五分之一的差距,发布相关资源供重复测量。
CAViAR:用于真实场景细粒度事故推理的因果视频数据集
机构 * NEC Laboratories, America(美国NEC实验室)
专题命中 代码评测 :repository(abstract,abstract_cn)
AI总结 该研究推出人工标注的真实事故视频基准CAViAR,测试发现现有VLMs存在感知-推理差距,无法可靠将驾驶场景主体行为映射到责任类别。
Comments Accepted to ECCV 2026 Workshop DriveX
仓颉基准:在低资源通用编程语言上评估大语言模型
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
AI总结 本文提出CangjieBench,用于评估大语言模型在低资源通用编程语言上的表现,通过248个高质量样本覆盖文本到代码和代码到代码任务,发现语法受限生成在准确性和计算成本之间取得最佳平衡。
Comments Accepted by ESEM 2026
ORCA-bench:语言模型智能体的值班待命准备程度如何?
机构 * Cornell Tech(康奈尔科技学院) ; Traversal ; Columbia University(哥伦比亚大学)
专题命中 代码评测 :coding agent(abstract);分类 cs.SE、cs.CL、cs.AI
AI总结 该研究推出ORCA-bench基准,评估编码智能体在生产级值班待命场景下的根本原因分析能力,发现现有前沿智能体表现不佳,凸显其距离安全胜任生产可靠性工作仍有较大差距。
CodexGraph:通过代码图数据库连接大型语言模型和代码库
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.CL、cs.AI
AI总结 研究旨在解决大型语言模型处理整个代码库的难题,提出通过将LLM代理与代码库图数据库接口集成的CodexGraph系统,利用图数据库特性实现精确上下文检索和代码导航,经多基准评估及实际应用开发,展现其在软件工程中的竞争力与潜力。
Comments work in progress
DFAH-Bench:金融决策中可观测智能体不稳定性的基准测试
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究金融决策中智能体行为稳定性,引入DFAH-Bench基准,通过多渠道衡量。发现仅结果一致性不能完全反映稳定性,前沿模型存在决策与工具路径一致性差距,识别出三种行为模式,并开源相关代码和数据。
Comments 15 pages, 3 figures. Code, sanitized replay logs, one-command reproduction (make reproduce-paper), and an interactive results explorer: https://github.com/ibm-client-engineering/output-drift-financial-llms
Git辅助工具:基于规划的Git仓库更新支持
机构 * AI Research, JPMorganChase(人工智能研究,摩根大通)
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.CL、cs.AI
AI总结 研究针对git工具对开发者有挑战的问题,提出结合大语言模型与自动规划的Git-Assistant,经合成和随机git环境评估,证明该方法能提升仓库管理可靠性、减少错误,展现混合人工智能方法在开发者辅助上的潜力。
Comments Pending permissions from the private company
ResearchClawBench: 端到端自主科学研究基准
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 代码评测 :coding agent(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出ResearchClawBench基准,包含10个领域40个任务,通过多模态评分标准评估自主科研能力,最强智能体仅得21.5分,揭示当前系统在实验协议、证据匹配和科学核心方面的不足。