arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13077cs.CYcs.AIcs.SE

单一文化的搭便车指南

The Hitchhiker's Guide to Monoculture: AI Homogenizes Syntax, Not (Necessarily) Semantics

Gordon Burtch

首次发表
浏览论文内容

中文总结 AI 辅助

研究通过Kaggle竞赛提交代码,探讨人工智能编码助手对代码的影响,采用多种方法在不同层面研究代码同质化,发现句法有显著同质化,语义同质化不明显,表明编码助手规范了实现细节但未使方法和策略趋同。

中文摘要 AI 辅助

大语言模型(LLMs)常产生同质化输出,引发对人工智能编码助手可能导致开发者创建的软件工件趋同的担忧。本文通过2019年至2026年年中的Kaggle竞赛提交代码来研究代码同质化。首先记录了向随机种子值42的广泛趋同。接着在两个聚合和抽象层面更广泛地研究同质化,在提交层面测量竞赛内提交代码的平均成对相似度,在竞赛层面测量提交代码的概念跨度,分别采用TF-IDF表示法和Voyage 3代码嵌入。结果表明在个体和集体层面都有显著的句法同质化,但语义同质化证据不足。这表明人工智能编码助手正在规范实现细节,但尚未使程序员采用的方法和解决问题策略趋同。

英文摘要

Large language models (LLMs) have been widely reported to homogenize human expression and thought. However, I argue that convergence in language need not imply convergence in ideas, and existing evidence rarely distinguishes between the two. I demonstrate this through a study of software development, where AI assistants have diffused fastest and where syntax can be readily separated from semantics (conceptual approach or intent). Using Kaggle contest submissions from 2019 to mid-2026, I first document convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding programming-culture convention associated with Douglas Adams' comedy novel The Hitchhiker's Guide to the Galaxy. I then measure homogenization in code syntax and approach more generally, quantifying within-contest code submission similarity using term-frequency-inverse-document-frequency (TF-IDF) n-gram representations, which capture surface syntax, and Voyage code-3 retrieval embeddings, which capture intent and approach. Submissions have become more alike in literal syntax, but I find no evidence of homogenization in approach or intent. I conclude that shared AI tools are homogenizing how Kaggle contestants express their solutions, not the problem-solving strategies they employ.

发表机构

  • Questrom School of Business, Boston University(波士顿大学奎斯特罗姆商学院)

机构由 AI 辅助整理,请以论文原文为准。

↑