GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
GBQA:用于评估LLM作为质量保证工程师的博弈基准
机构 * The University of Hong Kong(香港大学) ; Independent Researcher(独立研究者) ; Westlake University(西湖大学) ; Datawhale Org(Datawhale组织)
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI
AI总结 本文提出GBQA基准,通过30款游戏和124个经人类验证的bug测试LLM自主发现软件缺陷的能力,实验表明最佳模型仅能识别48.39%的bug。
Comments Accepted as a workshop paper at the Fourteenth International Conference on Learning Representations (ICLR 2026)