Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Preprint
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Preprint
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Accepted by ACL 2024 main conference
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.LG
Comments Website - https://livecodebench.github.io/
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI、cs.PL
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments NAACL 2024 Findings
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.AI
Comments Accepted to Findings of EACL 2024
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI、cs.LG
Comments Accepted by ICSE-SEIP 2024
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.AI、cs.LG
Comments 71 pages, 29 figures
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 12 pages, 4 figures, Oral Presentation at 3rd Workshop on Efficient Natural Language and Speech Processing (ENLSP-III), NeurIPS 2023
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 24 pages, 1 figure, 3 tables
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.LG
Comments Project site with code and data: https://intercode-benchmark.github.io
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published as workshop paper at Practical ML for Developing Countries Workshop @ ICLR 2020
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by ACL 2023 Findings. The first three authors contributed equally
专题命中 代码评测 :program synthesis(abstract);分类 cs.CL、cs.AI、cs.LG
Comments humaneval results, clarity
专题命中 代码评测 :code model(abstract);分类 cs.CL、cs.LG、cs.PL
Comments Accepted to ICML 2023, Code and data release: https://github.com/google-research/babelcode
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI、cs.PL
Comments In the proceedings of Advances in Neural Information Processing Systems, 2022
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.AI、cs.PL
Comments The 37th IEEE/ACM International Conference on Automated Software Engineering
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Title and relevant changes are made
专题命中 代码评测 :program repair(abstract);分类 cs.SE、cs.AI、cs.PL
Comments 3 pages, 3 tables, 1 GitHub url: https://github.com/pmorvalho/C-Pack-IPAs
专题命中 代码评测 :program synthesis(abstract);分类 cs.SE、cs.AI、cs.LG
Comments Accepted to the Technical Track of ICSE 2022
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by AAAI2021
机构 * Microsoft(微软公司) ; Hong Kong Baptist University(香港 Baptist 大学)
专题命中 代码评测 :code generation(abstract,comments);分类 cs.CL、cs.AI
Comments Large Language model, Code Generation, Code LLMs.This paper has been accepted to ICLR 2024. Please cite the ICLR version
Journal ref The Twelfth International Conference on Learning Representations (ICLR 2024)
专题命中 代码评测 :program repair(abstract,comments);分类 cs.SE、cs.AI
Comments Accepted to the 6th International Workshop on Automated Program Repair (APR 2025)
专题命中 代码评测 :code model(abstract,comments);分类 cs.SE、cs.LG
Comments Accepted to IEEE Transactions on Software Engineering. Extension of our previous paper "What do pre-trained code models know about code?" (ASE 2021, arXiv:2108.11308). 21 pages
通过TRIAD实现多跳RAG评估自动化:从上下文提取到验证数据集生成
机构 * University of Innsbruck(因斯布鲁克大学)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 本文提出TRIAD三阶段自动化多跳RAG评估数据集生成方法,经与MuSiQue、HotpotQA对比验证,该生成数据集可有效评估特定领域RAG系统性能,相关代码与验证结果已公开。
Comments Accepted at the 19th International Natural Language Generation Conference (INLG 2026)
星际争霸II中AlphaStar的AI控制之运行时动作干预
机构 * University of New South Wales(新南威尔士大学) ; CSIRO(澳大利亚联邦科学与工业研究组织)
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
AI总结 该研究提出运行时动作干预(RAI)机制,将其应用于AlphaStar复现版并开展《星际争霸II》人类实验,发现向用户披露AI对手能力会显著影响人类对其公平性、毒性的感知,需将执行栈控制与能力披露分开评估。
超越基准:基于拟人化与生命周期导向路线图的大语言模型评估
机构 * Department of Networks, China Mobile Communications Group Co.,Ltd.(中国移动通信集团有限公司网络部) ; Xidian University(西安电子科技大学) ; Oklahoma State University(俄克拉荷马州立大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 本研究针对LLM基准分数与实际效用脱节问题,建立诊断本体并提出含IQ、PQ、EQ、VQ的拟人化评估框架,分析200余个基准后为LLM开发提供战略指引。
Comments Preprint. Under Review
大型语言模型在软件工程与软件安全交叉领域:一项以证据为中心的结构化调查与研究议程
机构 * Nanjing Liancheng Intelligent Technology Group(南京连城智能科技集团)
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.AI
AI总结 本研究调查了LLMs在软件工程与软件安全交叉领域的进展,提出保证框架,识别有效性威胁与最低报告协议,制定研究议程,主张以任务适配证据而非单一基准评判模型能力。
TokEval:一个分词器评估套件
机构 * EPFL(洛桑联邦理工学院)
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.LG
AI总结 本研究推出TokEval框架,通过信息论、结构敏感等指标评估分词器,经实验验证其可预测下游模型性能,助力更原则性的分词器评估。
Comments Published as a conference paper at COLM 2026; Library hosted at https://github.com/cimeister/tokenizer-intrinsic-evals