arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

识别、分类和解释AI生成代码中偏见的框架

A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code

Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez

arXiv 2609.30642首次发表:更新:

发表机构

University of British Columbia, Kelowna; Federal University of Pará(不列颠哥伦比亚大学基洛纳分校; 帕拉联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一个基于分类法的框架,利用大型语言模型自动识别和解释AI生成代码中的偏见,实验显示模型在分类和解释上均与专家高度一致。

AI 中文摘要

随着大型语言模型(LLMs)被整合到软件开发工作流程中,人们对AI生成代码中无意偏见的担忧日益增加。尽管有证据表明这些偏见确实存在,但有限的研究系统地识别、分类和解释了它们。本研究调查了AI生成代码中的偏见,并评估了LLMs能否通过一个基于分类法的框架可靠地识别和解释偏见。我们扩展了一个现有的有偏见的AI生成Python代码数据集,并手动为代码片段标注了偏见类别和人工撰写的理由,以建立一个基准真值数据集。利用该数据集,我们通过上下文学习(ICL)评估了专有和开源LLMs作为自动化偏见检测和理由生成系统。最后,我们使用结构化理由和代码识别指标分析了LLM生成的解释与人工撰写的理由之间的相似性。我们的研究结果表明,LLMs能够有效支持代码偏见的识别和解释。Gemini实现了80.14%的分类准确率,精确率为84.0%,召回率为95.7%,而最佳开源替代方案Qwen3-coder实现了82.45%的准确率,68.64%的精确率和80.22%的召回率。此外,模型相对于人工推理的理由相似度得分分别为80.4%和80.14%,代码识别相似度得分分别为86.0%和87.82%。这些结果表明,LLMs能够检测生成的Python代码中的偏见逻辑,并产生与专家解释基本一致的解释。

英文摘要

As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and explained them. This study investigates bias in AI-generated code and evaluates whether LLMs can reliably identify and explain it through a taxonomy-driven framework. We extended an existing dataset of biased AI-generated Python code and manually annotated snippets with bias categories and human-authored justifications to establish a ground-truth dataset. Using this dataset, we evaluated proprietary and open-source LLMs as automated bias detection and justification systems through ICL. Finally, we analyzed similarity between LLM-generated explanations and human-authored justifications using structured justification and code identification metrics. Our findings demonstrate that LLMs can effectively support code bias identification and explanation. Gemini achieved 80.14% classification accuracy, with 84.0% precision and 95.7% recall, while the best open-source alternative, Qwen3-coder, achieved 82.45% accuracy, 68.64% precision, and 80.22% recall. Additionally, the models achieved justification similarity scores of 80.4% and 80.14%, respectively, relative to human-authored reasoning, and code identification similarity scores of 86.0% and 87.82%. These results suggest that LLMs can detect biased logic in generated Python code and produce explanations that substantially align with expert interpretations.

CommentsUnder Review at ACM TOSEM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑