发表机构
King’s College London; MATS Research(伦敦国王学院; MATS研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出LangChoiceBench基准,评估25种LLMs的代码语言选择,发现其普遍过度选择Python、一致性低,小开放模型偏好更强,还揭示了模型的虚假证据等失败模式。
AI 中文摘要
已有研究表明,大语言模型(LLMs)在生成项目级代码时表现出强烈的Python偏好,但目前尚无系统方法来衡量新模型的此类行为。为填补这一空白,我们推出LangChoiceBench——一款用于测量Python偏好、建议-实现一致性及语言多样性的项目级代码生成基准。LangChoiceBench涵盖7个软件领域的28个项目,其中Python通常并非合适的默认选择。我们评估了25种不同的LLMs,发现Python仍被过度选择,建议-实现一致性较低,且规模较小的开放权重模型通常表现出更强的Python偏好和更低的语言多样性。我们进一步分析了9826条推理轨迹,发现多数Python选择是自动做出的或主要由便捷性驱动,而非明确考虑项目需求;在少量但重要的案例中,模型会编造选择Python的上下文支持(我们称之为“虚假证据”的失败模式),或生成与自身推理中所选语言矛盾的代码。
英文摘要
Large language models (LLMs) have been shown to exhibit strong Python preferences when generating project-level code, but there is currently no systematic way to measure this behaviour across new models. To bridge this gap, we introduce LangChoiceBench, a project-level code-generation benchmark for measuring Python preference, recommendation-implementation consistency, and language diversity. LangChoiceBench covers 28 projects across seven software areas where Python is often a poor default. We evaluate 25 diverse LLMs and find that Python remains heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models generally show stronger Python preference and lower language diversity. We further analyse 9,826 reasoning traces and find that most Python choices are automatic or driven primarily by ease, rather than explicit consideration of project requirements. In a smaller but important set of cases, models fabricate contextual support for choosing Python - a failure mode we call phantom evidence - or produce code that contradicts the language selected in their own reasoning.
Comments19 pages, 9 tables, 2 figures