ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
机构 * Tencent Hunyuan Team(腾讯文脉团队)
专题命中 代码评测 :code generation(title,abstract);分类 cs.SE、cs.CL
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
机构 * Tencent Hunyuan Team(腾讯文脉团队)
专题命中 代码评测 :code generation(title,abstract);分类 cs.SE、cs.CL
专题命中 代码评测 :program synthesis(title);分类 cs.SE、cs.LG、cs.PL
Comments 25 pages, 1 figure; data available at https://github.com/Beneficial-AI-Foundation/vericoding-benchmark
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments To appear in NeurIPS 2025. Welcome your submission to challenge our leaderboard at: https://code4db.github.io/parrot-bench/. Also visit our code repository at: https://github.com/weAIDB/PARROT
专题命中 代码评测 :repository(abstract)
机构 * School of Information and Communication Engineering, University of Electronic Science and Technology of China(信息与通信工程学院,电子科学与技术大学) ; Tianjin Key Laboratory of Visual Computing and Intelligent Perception, School of Computer Science, Nankai University(视觉计算与智能感知天津重点实验室,计算机科学学院,南开大学) ; Faculty of Mathematics and Computer Science, Weizmann Institute of Science(数学与计算机科学学院,魏兹曼科学研究所)
专题命中 代码评测 :repository(abstract)
Comments Accepted by Proceedings of the IEEE