AI 中文总结
该研究首次系统探究基于推理的代码生成中的社会偏见,发现推理可降低偏见但会影响代码质量,进而提出ProbeDebias框架,有效检测并缓解代码偏见,同时保持较高质量。
AI 中文摘要
大语言模型(LLMs)越来越多地被用于代码生成,但生成的程序可能会通过对敏感人口统计属性的不公平或差别对待表现出社会偏见。尽管先前的工作主要研究直接代码生成,但基于推理的生成中的偏见仍未得到充分探索。我们对基于推理的代码生成中的社会偏见开展了首次系统研究,在三个以人类为中心的决策场景中的现实偏见敏感任务上,评估了9种标准LLM和大型推理模型(LRMs)。我们发现,推理通常会降低偏见,将平均偏见率从0.64降至0.40,但该效果在不同模型间差异显著;同时代码质量并未始终保持,平均质量从0.72降至0.59。有偏见的推理强烈预示着有偏见的代码,仅调整生成配置不足以实现稳健缓解。基于这些发现,我们提出ProbeDebias,这是一个感知推理的框架,可在代码生成前检测并重写有偏见的推理轨迹。ProbeDebias在推理偏见检测上达到87.76%的F1值,平均降低代码偏见83.73%,同时在很大程度上保持了质量;与SOTA基线相比,其进一步将平均偏见降低52.70%-54.42%,并将质量提升9.79%-36.79%。这些结果凸显了推理阶段分析对于可信代码生成的价值。
英文摘要
Large language models (LLMs) are increasingly used for code generation, yet generated programs may exhibit social bias through unfair or differential treatment of sensitive demographic attributes. While prior work mainly studies direct code generation, bias in reasoning-based generation remains underexplored. We conduct the first systematic study of social bias in reasoning-based code generation, evaluating 9 standard LLMs and large reasoning models (LRMs) on realistic bias-sensitive tasks across three human-centered decision scenarios. We find that reasoning generally reduces bias, lowering the average bias rate from 0.64 to 0.40, but the effect varies substantially across models. Meanwhile, code quality is not consistently preserved, with the average quality dropping from 0.72 to 0.59. Biased reasoning strongly predicts biased code, and adjusting generation configurations alone is insufficient for robust mitigation. Based on these findings, we propose ProbeDebias, a reasoning-aware framework that detects and rewrites biased reasoning traces before code generation. ProbeDebias achieves 87.76% F1 for reasoning-bias detection and reduces code bias by 83.73% on average while largely preserving quality. Compared with SOTA baselines, it further reduces average bias by 52.70%-54.42% and improves quality by 9.79%-36.79%. These results highlight the value of reasoning-stage analysis for trustworthy code generation.
CommentsASE 2026