Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
在软件工程中评估AI模型:一项综述、搜索工具和统一方法以提高基准质量
机构 * EEMCS faculty, Delft University of Technology(代尔夫特理工大学EEMCS学院)
专题命中 代码评测 :code generation(abstract);code model(abstract);分类 cs.SE、cs.AI
AI总结 本文提出BenchScout和BenchFrame,通过统一框架提升软件工程AI模型的基准质量,改进现有基准并验证其有效性。
Comments Accepted for publication in IEEE Transactions on Software Engineering