arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02772cs.SE

从代码库到语言模型:Linux 基金会存储库中的非包容性命名

From Codebases to LLMs: Non-Inclusive Naming in Linux Foundation Repositories

Honghao Tan, Md Nafiu Rahman, Shin Hwei Tan

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用 NISCAN 框架检测 Linux 基金会存储库中代码及相关工件的非包容性术语,对 461 个存储库进行生态系统规模研究,分析术语变化及影响因素,还通过案例研究评估大语言模型处理情况。

中文摘要 AI 辅助

自2020年以来,Linux 基金会和多组织包容性命名倡议鼓励开源项目替换非包容性术语。本文提出 NISCAN 框架检测非包容性术语,对461个存储库研究发现非包容性术语自2020年下降约47%但仍未完全采用,还研究了对人工智能辅助软件开发的影响。

英文摘要

Since 2020, the Linux Foundation and the multi-organization Inclusive Naming Initiative (INI) have encouraged open-source projects to replace non-inclusive terms such as master/slave and whitelist/blacklist. Although these recommendations have been widely adopted, there is limited empirical evidence on their long-term adoption across Linux Foundation (LF) projects or their implications for AI-assisted software development. In this paper, we present NISCAN, a multilingual static-analysis framework that detects non-inclusive terminology across source code and related software artifacts using the INI vocabulary. Using NISCAN, we conduct the first ecosystem-scale study of inclusive naming across 461 Linux Foundation repositories. Our analysis shows that non-inclusive terminology has declined by approximately 47% since 2020, yet adoption remains incomplete: 62.7% of repositories still contain at least one Tier-1 non-inclusive identifier, while most remaining terminology resides outside source code in documentation, comments, configuration files, and other software artifacts. We further show that repository size, programming language, project functionality, and ecosystem are stronger predictors of term inclusiveness in LF repositories rather than foundation governance. To examine the implications for AI-assisted software development, we conduct a case study evaluating whether large language models (LLMs) can reconstruct legacy non-inclusive identifiers from surrounding program context. The results show that historical naming decisions remain embedded in model predictions even after identifiers have been renamed. Overall, our study findings provide the first ecosystem-scale assessment of inclusive naming adoption within the Linux Foundation and highlight the importance of addressing terminology residue to support responsible naming and ethically sourced code generation.

↑