arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22511cs.SE

身份还是提示噪声?大语言模型代码生成的校准不变性审计

Identity or Prompt Noise? A Calibrated Invariance Audit of LLM Code Generation

Maksim E. Eren, Ryan Barron, Eric Michalak, Charles Nicholas, Manish Bhattarai

首次发表
浏览论文内容

中文总结 AI 辅助

本审计通过3070万次生成实验发现,职业人设对代码形式有微小但可复现的影响,而国家与性别影响不显著,表明身份线索不导致稳定危害。

中文摘要 AI 辅助

身份线索与固定的编程规范无关,但原始的反事实差异可能源于样本不均衡和提示措辞。我们对HumanEval+和MBPP+上七个检查点的3070万次执行的Python生成结果进行了审计,并辅以一个探索性的550B切片,在模型指定的性别、国家和职业人设下进行。任务内随机化和错误发现率控制识别出职业是最一致的结构性轴:CodeBLEU离散度在14个模型-基准组合单元中的10个中超过其可交换性零假设,在8/12个全覆盖单元中仍然显著,并且在每个配对单元中都超过国家比率,尽管中位数超额仅为0.141分。在六个高通过率单元中,职业离散度在令牌相似性、长度、注释、参考相似性和复杂性方面均能复制,而通过率离散度在任何一个单元中都不显著。国家在14个单元中的12个中主导原始离散度,但其中位数校准比率为1.00。性别相关的变异在此设计中无法与人设措辞区分开来。因此,证据支持代码形式存在微小、可复现的职业条件变化,而非对指定身份的稳定不利或已证明的下游危害。我们利用这一发现来推动对代码生成中偏见潜在影响的更广泛评论,包括模型可能根据先前对话上下文中可用的身份信息来调整其输出的可能性。

英文摘要

Identity cues are irrelevant to a fixed programming specification, but raw counterfactual differences can arise from unequal samples and prompt wording. We audit 30.73 million executed Python generations from seven checkpoints on HumanEval+ and MBPP+, supplemented by an exploratory 550B slice, under model-assigned gender, country, and occupation personas. Within-task randomization and false-discovery-rate control identify occupation as the most consistent structural axis: CodeBLEU dispersion exceeds its exchangeability null in 10/14 model--benchmark cells, remains significant in 8/12 full-coverage cells, and exceeds the country ratio in every paired cell, although the median excess is only 0.141 points. In six high-pass-rate cells, occupation dispersion replicates across token similarity, length, comments, reference similarity, and complexity, while pass-rate dispersion is significant in none. Country leads raw dispersion in 12/14 cells but has a median calibrated ratio of 1.00. Gender-associated variation cannot be separated from persona wording in this design. Thus, the evidence supports small, reproducible occupation-conditioned changes in code form, not stable disadvantage to named identities or demonstrated downstream harm. We use this finding to motivate a broader commentary on the potential impacts of bias in code generation, including the possibility that models may condition their outputs on identity information available from prior conversational context.

发表机构

  • LANL(洛斯阿拉莫斯国家实验室)
  • UMBC(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

↑