当你的身份能改变你获得的代码:LLM代码生成中人格诱导偏差的研究
When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation
浏览论文内容
中文总结 AI 辅助
本研究通过大规模实证发现,用户人口统计信息(人格)会显著影响LLM代码生成的质量,导致推理和输出中出现人口统计标记泄漏,并引起正确性等指标的可测量变化,凸显了LLM辅助开发中的新风险。
中文摘要 AI 辅助
大型语言模型(LLM)被广泛用作编程助手,但用户的人口统计信息是否以及如何影响生成代码的技术质量仍不清楚。我们针对基于LLM的代码生成中的人格诱导偏差进行了一项大规模实证研究,重点关注一个专有模型(Gemini 2.5 Pro)和一个开放权重模型(GPT-OSS-120B)。使用涵盖国籍、性别和经验水平的18种人口统计人格,我们将人格诱导提示与中性基线进行比较。在35,000多个生成的程序中,我们分析了推理和响应中的人口统计标记泄漏,以及功能正确性、可维护性、代码风格和安全性方面的差异。我们的结果表明,人口统计线索经常反映在LLM的推理和输出中。尽管人口统计标记在语义上与任务无关,但它们出现在高达65%的响应和70%的推理轨迹中。在LiveCodeBench上,人格提示与Gemini模型的正确性得分平均降低1.54个百分点相关,其中一种人格的得分下降了3.6%(优势比=0.51)。相比之下,GPT-OSS模型在所有人格上的准确率提高了3.4%至5.7%(优势比=1.8至3.0)。可维护性和代码风格指标显示出统计显著但效应量可忽略不计的差异(所有Cliff's δ < 0.15),安全漏洞则没有表现出系统性的人格特定模式。总体而言,我们的结果表明,即使在纯技术任务中,用户人口统计信息的存在也与LLM推理和代码质量的可测量变化相关,并且这些效应在多个模型中均存在。我们的工作凸显了LLM辅助软件开发中一个未被充分审视的风险。
英文摘要
Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemini 2.5 Pro) and an open-weight model (GPT-OSS-120B). Using 18 demographic personas spanning nationality, gender, and experience level, we compare persona-induced prompts against a neutral baseline. Across 35,000+ generated programs, we analyze demographic marker leakage in reasoning and responses, as well as differences in functional correctness, maintainability, code style, and security. Our results show that demographic cues are frequently reflected in LLM reasoning and outputs. Demographic markers appear in up to 65% of responses and 70% of reasoning traces, despite being semantically irrelevant to the tasks. On LiveCodeBench, persona prompting were associated with lower correctness scores of the Gemini model by an average of 1.54 percentage points, with one persona exhibiting a decrease of 3.6% (odds ratio = 0.51). In contrast, the accuracy of the GPT-OSS model improved by 3.4 - 5.7% across all personas (odds ratios = 1.8 - 3.0). Maintainability and code style metrics show statistically significant but negligible effect sizes (all Cliff's δ < 0.15), and security vulnerabilities exhibit no systematic persona-specific patterns. Overall, our results show that the presence of demographic information about users is associated with measurable variation in LLM reasoning and code quality even in purely technical tasks, and that these effects hold across models. Our work highlights an under-examined risk in LLM-assisted software development.
发表机构
- Extended author information available on the last page of the article
机构由 AI 辅助整理,请以论文原文为准。