arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

成本高效大型语言模型在算法编程任务上的实证评估

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

Chandimal Adikari, Nandika Herath

arXiv 2609.18052首次发表:更新:

发表机构

University of Wollongong; Western Sydney University(伍伦贡大学; 西悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究实证评估成本高效LLM生成企业代码的可靠性,发现结构符合性高但计算正确性低,且可靠性与正确性呈反比关系。

AI 中文摘要

本研究实证评估了成本高效的大型语言模型(LLMs)是否能够被信任,依据书面规范生成企业代码。三个模型(Gemini Flash 3、GPT-5.4 mini 和 Claude Haiku 4.5)被要求将992个算法问题作为Java Spring Boot服务方法解决,并遵循规定的签名和数据传输对象规范,交叉四种模型与智能体编码工具组合及两种提示变体,形成八种配置,禁止迭代且明确禁止硬编码答案。八个问题陈述被保留,以探究模型对缺失输入的反应。产生的7,593个方法按一个八类结果分类法进行分类,该分类法描述每个方法在生成答案方面的行为,随后部署并执行,得到7,936次测量请求并与分类关联。结构符合性接近上限,但38.4%的方法并未计算其返回的值,且仅12.9%的返回答案是正确的。按结果类别进行条件分析显示,响应可靠性与正确性呈反比关系,而真正进行计算的方法回答频率最低,正确率为19.3%。局限性包括每种配置仅单次生成运行、部分测试平台覆盖、单次计时、语法分类以及可能的语料库污染。

英文摘要

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a mandated signature and data-transfer-object specification, crossing four model and agentic coding tool combinations with two prompt variants to yield eight configurations, with iteration forbidden and hardcoded answers explicitly prohibited. Eight problem statements were withheld to probe how models respond to missing input. The 7,593 resulting methods were classified by an eight-class outcome taxonomy describing what each does about producing an answer, then deployed and executed, giving 7,936 measured requests joined to that classification. Structural conformance approached ceiling, yet 38.4% of methods do not compute the value they returned and only 12.9% of returned answers were correct. Conditioning on outcome class shows that response reliability and correctness are inversely related, whereas genuinely computing methods answered least often and were correct 19.3%. Limitations include single generation runs per configuration, partial harness coverage, single-pass timing, syntactic classification, and probable corpus contamination.

Comments7 pages, 7 figures. Accepted for publication in the IEEE Proceedings of the 2026 International Conference on Advanced Computing Technologies (ICACT)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑